system
The system uses AI to analyze and generate personalized assembly instructions for educational blocks, addressing the lack of user creativity in conventional methods by providing novel and tailored assembly solutions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional educational block assembly instructions are fixed and do not fully elicit user creativity.
A system comprising a shooting unit, analysis unit, and generation unit that uses AI to analyze photographs of educational blocks and generate original assembly instructions tailored to the user's wishes, considering block characteristics, user history, and current projects.
The system enhances user creativity by generating personalized and novel assembly instructions, allowing users to gain new ideas and effectively reuse educational blocks.
Smart Images

Figure 2026073272000001_ABST
Abstract
Description
Technical Field
[0004] ,
[0006] , , ,
[0005] , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that the assembly instructions of the educational blocks are fixed and the creativity of the user cannot be fully elicited.
[0005] The system according to the embodiment aims to elicit the creativity of the user and generate an original assembly instruction.
Means for Solving the Problems
[0006] The system according to this embodiment comprises a shooting unit, an analysis unit, a generation unit, and an input unit. The shooting unit takes a photograph of the educational building blocks. The analysis unit analyzes the photograph taken by the shooting unit. The generation unit generates an original assembly instruction manual based on the results of the analysis by the analysis unit. The input unit takes input of what the user wants to build. [Effects of the Invention]
[0007] The system according to this embodiment can draw out the user's creativity and generate original assembly instructions. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The educational block assembly support system according to an embodiment of the present invention is a system in which a generating AI generates original block assembly instructions for educational blocks taken from a photograph or a photograph taken on the spot. The educational block assembly support system works by having the user take a photograph of the educational blocks, and the generating AI analyzes the photograph to generate original assembly instructions. These instructions are generated to suit what the user wants to build (for example, a dinosaur or a castle). Furthermore, even if multiple different blocks are mixed and photographed, the system can generate a novel assembly method that incorporates all of them. This allows the user to gain new ideas and reuse the blocks. For example, the educational block assembly support system works by having the user take a photograph of the educational blocks. Next, the generating AI analyzes the photograph and generates original assembly instructions. These instructions are generated to suit what the user wants to build (for example, a dinosaur or a castle). Furthermore, even if multiple different blocks are mixed and photographed, the system can generate a novel assembly method that incorporates all of them. This allows the user to gain new ideas and reuse the blocks. Thus, the educational block assembly support system allows the user to gain new ideas and reuse the blocks.
[0029] The educational block assembly support system according to this embodiment comprises a shooting unit, an analysis unit, a generation unit, and an input unit. The shooting unit takes photographs of the educational blocks. The shooting unit can take photographs of the educational blocks using, for example, a smartphone or a digital camera. The shooting unit can also automatically optimize the arrangement and angle of the blocks when taking photographs. Furthermore, the shooting unit can automatically identify the color and shape of the blocks and select the optimal shooting settings. The analysis unit analyzes the photographs taken by the shooting unit using a generation AI. The analysis unit can analyze the types and arrangement of the blocks using, for example, an image recognition algorithm. Furthermore, the analysis unit can analyze photographs taken with multiple different blocks mixed together and generate novel assembly methods that combine them all. Furthermore, the analysis unit can optimize the analysis results by considering, for example, the material and texture of the blocks. The generation unit uses a generation AI to generate original assembly instructions based on the results analyzed by the analysis unit. The generation unit can generate assembly instructions tailored to what the user wants to make. Furthermore, the generation unit can suggest the optimal assembly method by referring to the user's past assembly history. Furthermore, the generation unit can customize the assembly instructions based on the user's current projects or areas of interest, for example. The input unit allows the user to input what they want to create. The input unit can input the user's wishes using, for example, text input or voice input. The input unit can also estimate the user's emotions and adjust the display method of the input interface based on the estimated emotions. Furthermore, the input unit can suggest the optimal input method by referring to the user's past input history, for example. As a result, the educational block assembly support system according to the embodiment allows the user to gain new ideas and reuse the blocks.
[0030] The photography unit takes pictures of the educational building blocks. The photography unit can take pictures of the blocks using, for example, a smartphone or digital camera. Specifically, it launches the camera app on a smartphone and takes pictures while checking the arrangement of the blocks. For digital cameras, it is recommended to use a tripod for stable shooting. The photography unit can also automatically optimize the arrangement and angle of the blocks before shooting. This utilizes the camera's autofocus function and AI-based image correction technology. Furthermore, the photography unit can automatically identify the color and shape of the blocks and select the optimal shooting settings. This allows the photography unit to accurately capture the detailed characteristics of the blocks and provide high-quality image data to the analysis unit. For example, it can automatically adjust the camera's white balance and exposure settings to ensure that the block colors are accurately reproduced. Also, even when multiple blocks are mixed together, the photography unit optimizes the position and angle of each block during shooting, enabling the analysis unit to process the data efficiently. This allows the photography unit to support users in easily taking high-quality photos and improve the overall accuracy and efficiency of the system.
[0031] The analysis unit uses generative AI to analyze photographs taken by the photography unit. For example, the analysis unit can analyze the type and arrangement of blocks using image recognition algorithms. Specifically, it uses a deep learning-based image recognition model to identify the shape, color, and size of the blocks. For instance, it can use a convolutional neural network (CNN) to analyze the edges and textures of the blocks and extract the characteristics of each block. The analysis unit can also analyze photographs of mixed blocks and generate novel assembly methods using all of them. This involves a process where the generative AI learns the block arrangement patterns and proposes new assembly methods. Furthermore, the analysis unit can optimize the analysis results by considering the material and texture of the blocks. For example, it can identify whether a block is made of plastic or wood and propose an assembly method accordingly. This allows the analysis unit to provide the optimal assembly method based on the type and characteristics of the blocks the user possesses. Additionally, the analysis unit can learn user preferences and tendencies based on past analysis data to provide more personalized analysis results. This enables the analysis unit to perform flexible analysis tailored to user needs and improve the overall user experience of the system.
[0032] The generation unit uses a generation AI to generate original assembly instructions based on the results analyzed by the analysis unit. For example, the generation unit can generate assembly instructions tailored to what the user wants to create. Specifically, when the user inputs the type or theme of their desired creation through the input unit, the generation AI generates the optimal assembly procedure based on that information. For example, if the user inputs that they want to build a "robot," the generation unit will generate a detailed instruction manual for assembling the robot based on the block types and placement information obtained from the analysis unit. The generation unit can also suggest the optimal assembly method by referring to the user's past assembly history. This includes saving data on previously created creations and customizing new assembly procedures based on that data. Furthermore, the generation unit can customize the assembly instructions based on the user's current projects or areas of interest. For example, if the user wants to create a "space" themed creation, the generation unit will generate instructions incorporating space-related block placements and designs. In this way, the generation unit provides assembly instructions tailored to the user's individual needs and interests, supporting them in enjoying the block building process more. Additionally, the generation unit can save the generated instructions in PDF or image format, allowing the user to access them at any time. This allows the generation unit to provide users with high-quality, customized assembly instructions, improving the overall convenience and satisfaction of the system.
[0033] The input section is where the user enters what they want to create. The input section allows users to input their wishes using methods such as text input or voice input. Specifically, users can input text using a smartphone or tablet keyboard, or communicate their wishes by voice using voice recognition technology. For example, if a user voice-inputs "I want to make a dinosaur," the voice recognition system converts that into text and transmits it to the system. Furthermore, the input section can estimate the user's emotions and adjust the display of the input interface based on those emotions. This includes a process of estimating emotions from the user's facial expressions and tone of voice using facial recognition and voice analysis technologies. For example, if the user appears happy, the interface's colors and design will be brightened; conversely, if they appear tired, it will be changed to a simpler, more calming design. Additionally, the input section can suggest the optimal input method by referring to the user's past input history. This includes a function that predicts frequently used phrases and keywords based on data the user has previously entered, assisting with input. For example, if a user has previously entered "I want to make a car," the keyword "car" will be automatically suggested the next time they enter a similar wish. This allows the input section to support users in easily and quickly entering their preferences, improving the overall usability of the system.
[0034] The generation unit can generate assembly instructions tailored to what the user wants to build. For example, if the user wants to build a dinosaur, the generation unit will generate assembly instructions for a dinosaur. It can also generate assembly instructions for a castle if the user wants to build one, and for example, if the user wants to build a car. This improves user satisfaction by generating assembly instructions that match the user's wishes. Some or all of the above processing in the generation unit is performed using a generative AI. For example, the generation unit can use a generative AI model that takes what the user wants to build as input and generates assembly instructions based on that input.
[0035] The analysis unit includes a mixed analysis unit that analyzes photographs taken with multiple different blocks mixed together and generates novel assembly methods using all of them. For example, the analysis unit can analyze photographs taken with blocks from different manufacturers mixed together. It can also analyze photographs taken with blocks of different shapes mixed together. Furthermore, it can analyze photographs taken with blocks of different colors mixed together. This allows for the generation of novel assembly methods using multiple different blocks, thereby providing new ideas. Some or all of the above processing in the analysis unit is performed using generative AI. For example, the analysis unit can use a generative AI model that takes a photograph taken with multiple different blocks mixed together as input and generates a novel assembly method based on that photograph.
[0036] The camera unit can automatically optimize the placement and angle of blocks during shooting. For example, the camera unit can use AI to analyze the block placement in real time and shoot at the optimal angle. The camera unit can also use AI to automatically adjust the block placement to capture the most aesthetically pleasing composition. Furthermore, the camera unit can use AI to consider the position of the block shadows and shoot with optimal lighting. This allows for beautiful compositions by optimizing the block placement and angle. Some or all of the above processes in the camera unit are performed using AI. For example, the camera unit can use an AI model that receives block placement data as input and calculates the optimal angle based on that data.
[0037] The shooting unit can automatically identify the color and shape of the blocks during shooting and select the optimal shooting settings. For example, the shooting unit can use AI to identify the color of the blocks, set the optimal white balance, and take a picture. It can also use AI to identify the shape of the blocks, select the optimal focus settings, and take a picture. Furthermore, it can use AI to simultaneously identify the color and shape of the blocks, select the optimal exposure settings, and take a picture. This allows for the selection of optimal shooting settings according to the color and shape of the blocks, resulting in better photographs. Some or all of the above processing in the shooting unit is performed using AI. For example, the shooting unit can use an AI model that receives block color and shape data as input and calculates the optimal shooting settings based on that data.
[0038] The camera unit can suggest the optimal shooting method by referring to the user's past shooting history during shooting. For example, the camera unit's AI can analyze the user's past shooting history and suggest the most successful shooting method. The camera unit can also suggest the optimal shooting angle based on the user's past shooting history using AI. Furthermore, the camera unit can automatically select the optimal shooting settings by referring to the user's past shooting history using AI. In this way, the camera unit can suggest the optimal shooting method by referring to past shooting history. Some or all of the above processes in the camera unit are performed using AI. For example, the camera unit can use an AI model that receives the user's past shooting data as input and suggests the optimal shooting method based on that data.
[0039] The shooting unit can select the optimal shooting settings by considering the user's surrounding environment during shooting. For example, the shooting unit can use AI to measure the ambient light intensity and select the optimal exposure setting. It can also use AI to measure the ambient sound level and select the optimal shutter speed. Furthermore, it can use AI to measure the ambient temperature and select the optimal white balance. This allows for the selection of optimal shooting settings by considering the surrounding environment. Some or all of the above processing in the shooting unit is performed using AI. For example, the shooting unit can use an AI model that receives ambient environmental data as input and calculates the optimal shooting settings based on that data.
[0040] The analysis unit can optimize the analysis results by considering the material and texture of the block during the analysis. For example, the analysis unit's generating AI can identify the material of the block and select the optimal analysis algorithm. The analysis unit can also, for example, have the generating AI analyze the texture of the block and set the optimal analysis parameters. Furthermore, the analysis unit can, for example, have the generating AI consider both the material and texture of the block simultaneously to provide the optimal analysis results. This allows for more accurate analysis results by considering the material and texture of the block. Some or all of the above processes in the analysis unit are performed using the generating AI. For example, the analysis unit can use a generating AI model that receives block material and texture data as input and provides the optimal analysis results based on that data.
[0041] The analysis unit can automatically generate block combination patterns during analysis and propose the optimal combination. For example, the analysis unit's generating AI can generate the optimal combination pattern based on the shape and color of the blocks. The analysis unit can also propose the optimal combination pattern by considering the arrangement of the blocks, for example. Furthermore, the analysis unit can provide the optimal combination pattern by analyzing the frequency of block usage, for example. In this way, the analysis unit can propose the optimal combination by automatically generating block combination patterns. Some or all of the above processes in the analysis unit are performed using the generating AI. For example, the analysis unit can use a generating AI model that receives block shape and color data as input and generates the optimal combination pattern based on that data.
[0042] The analysis unit can select the optimal analysis method by referring to the user's past assembly history during analysis. For example, the analysis unit's generating AI can analyze the user's past assembly history and select the optimal analysis algorithm. The analysis unit can also, for example, have the generating AI set the optimal analysis parameters based on the user's past assembly history. Furthermore, the analysis unit can, for example, have the generating AI refer to the user's past assembly history and provide the optimal analysis results. In this way, the optimal analysis method can be selected by referring to the past assembly history. Some or all of the above processes in the analysis unit are performed using the generating AI. For example, the analysis unit can use a generating AI model that receives the user's past assembly data as input and selects the optimal analysis method based on that data.
[0043] The analysis unit can customize the analysis results based on the user's current projects and areas of interest during the analysis process. For example, the analysis unit can use a generative AI to analyze the user's current projects and provide optimal analysis results. Alternatively, the analysis unit can use a generative AI to consider the user's areas of interest and provide customized analysis results. Furthermore, the analysis unit can use a generative AI to provide optimal analysis results based on the user's project progress. This allows for the provision of more appropriate analysis results by customizing them based on the current projects and areas of interest. Some or all of the above-described processes in the analysis unit are performed using generative AI. For example, the analysis unit can use a generative AI model that receives the user's project data as input and customizes the analysis results based on that data.
[0044] The generation unit can propose the optimal assembly method by referring to the user's past assembly history during generation. For example, the generation unit's generation AI can analyze the user's past assembly history and propose the optimal assembly method. The generation unit can also provide the optimal assembly procedure based on the user's past assembly history, for example. Furthermore, the generation unit can propose the optimal assembly pattern by referring to the user's past assembly history, for example. In this way, the optimal assembly method can be proposed by referring to past assembly history. Some or all of the above processing in the generation unit is performed using the generation AI. For example, the generation unit can use a generation AI model that receives the user's past assembly data as input and proposes the optimal assembly method based on that data.
[0045] The generation unit can customize the assembly instructions based on the user's current projects and areas of interest during the generation process. For example, the generation unit's generation AI can analyze the user's current projects and provide customized assembly instructions. Alternatively, the generation unit can use its generation AI to consider the user's areas of interest and provide optimal assembly instructions. Furthermore, the generation unit can use its generation AI to provide customized assembly instructions based on the user's project progress. This allows for the provision of more appropriate assembly instructions by customizing them based on the current projects and areas of interest. Some or all of the above-described processes in the generation unit are performed using generation AI. For example, the generation unit can use a generation AI model that receives user project data as input and customizes the assembly instructions based on that data.
[0046] The generation unit can propose an optimal assembly method while considering the user's geographical location information during generation. For example, the generation unit's generation AI can propose an optimal assembly method based on the user's geographical location information. Furthermore, the generation unit can also provide region-specific assembly methods by considering the user's geographical location information. Additionally, the generation unit can propose an optimal assembly procedure by referencing the user's geographical location information. This allows for the provision of region-specific assembly methods by considering geographical location information. Some or all of the above-described processes in the generation unit are performed using the generation AI. For example, the generation unit can use a generation AI model that receives the user's geographical location data as input and proposes an optimal assembly method based on that data.
[0047] The generation unit can analyze the user's social media activity during generation and propose relevant assembly methods. For example, the generation unit's generation AI can analyze the user's social media activity and propose the optimal assembly method. The generation unit can also provide relevant assembly steps based on the user's social media activity, for example. Furthermore, the generation unit can refer to the user's social media activity and propose the optimal assembly pattern, for example. This allows for the provision of assembly methods based on the user's interests by analyzing social media activity. Some or all of the above processing in the generation unit is performed using a generation AI. For example, the generation unit can use a generation AI model that receives the user's social media data as input and proposes the optimal assembly method based on that data.
[0048] The input unit can suggest the optimal input method by referring to the user's past input history during input. For example, the input unit's generative AI can analyze the user's past input history and suggest the optimal input method. The input unit can also provide the optimal input procedure based on the user's past input history, for example. Furthermore, the input unit can suggest the optimal input pattern by referring to the user's past input history, for example. In this way, the optimal input method can be suggested by referring to past input history. Some or all of the above processing in the input unit is performed using generative AI. For example, the input unit can use a generative AI model that receives the user's past input data as input and suggests the optimal input method based on that data.
[0049] The input unit can customize the input content based on the user's current projects and areas of interest during input. For example, the input unit can use a generative AI to analyze the user's current projects and provide customized input content. Alternatively, the input unit can use a generative AI to consider the user's areas of interest and provide optimal input content. Furthermore, the input unit can use a generative AI to provide customized input content based on the user's project progress. This allows for the provision of more appropriate input content by customizing it based on the current projects and areas of interest. Some or all of the above processing in the input unit is performed using generative AI. For example, the input unit can use a generative AI model that receives the user's project data as input and customizes the input content based on that data.
[0050] The input unit can propose the optimal input method while considering the user's device information. For example, the input unit's generative AI can propose the optimal input method based on the user's device information. Furthermore, the input unit can, for example, have the generative AI consider the user's device information and provide an input procedure optimized for the device. Additionally, the input unit can, for example, have the generative AI refer to the user's device information and propose the optimal input pattern. This allows for the proposal of the optimal input method by considering the device information. Some or all of the above processing in the input unit is performed using generative AI. For example, the input unit can use a generative AI model that receives the user's device information as input and proposes the optimal input method based on that data.
[0051] The input unit can analyze the user's social media activity during input and suggest relevant input content. For example, the input unit can use a generative AI to analyze the user's social media activity and suggest the optimal input content. The input unit can also use a generative AI to provide relevant input steps based on the user's social media activity. Furthermore, the input unit can use a generative AI to refer to the user's social media activity and suggest the optimal input pattern. This allows for the provision of input content based on the user's interests by analyzing social media activity. Some or all of the above processing in the input unit is performed using a generative AI. For example, the input unit can use a generative AI model that receives the user's social media data as input and suggests the optimal input content based on that data.
[0052] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0053] The educational building block assembly support system can refer to the user's past assembly history and suggest new assembly methods based on the user's previously created works. For example, if the user previously created a dinosaur, the system can suggest a new variation of that dinosaur. Also, if the user previously created a castle, the system can suggest a new castle design based on that castle. Furthermore, if the user previously created a car, the system can suggest a new car model based on that car. In this way, the system can leverage the user's past assembly history to provide new ideas. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that takes the user's past assembly data as input and suggests new assembly methods based on that data.
[0054] The educational building block assembly support system can suggest region-specific assembly methods by considering the user's geographical location. For example, if the user is in Japan, the system can suggest assembly methods based on traditional Japanese architecture and culture. If the user is in the United States, the system can suggest assembly methods based on historical American buildings and culture. Furthermore, if the user is in Europe, the system can suggest assembly methods based on European architecture such as castles and churches. In this way, by considering geographical location, the system can provide region-specific assembly methods. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that takes the user's geographical location data as input and suggests the optimal assembly method based on that data.
[0055] The educational building block assembly support system can analyze a user's social media activity and suggest relevant building methods. For example, if a user frequently posts about dinosaurs on social media, the system can suggest how to build a dinosaur. Similarly, if a user frequently posts about castles, the system can suggest how to build a castle. Furthermore, if a user frequently posts about cars, the system can suggest how to build a car. In this way, by analyzing social media activity, the system can provide building methods based on the user's interests. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that takes the user's social media data as input and suggests the optimal building method based on that data.
[0056] The educational building block assembly support system can suggest the optimal input method by considering the user's device information. For example, if the user is using a smartphone, the system can suggest an input method optimized for touchscreens. If the user is using a tablet, the system can also suggest an input method optimized for large screens. Furthermore, if the user is using a desktop computer, the system can suggest an input method optimized for keyboards and mice. In this way, the system can suggest the optimal input method by considering device information. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that receives the user's device information as input and suggests the optimal input method based on that data.
[0057] The educational block assembly support system can suggest the optimal input method by referring to the user's past input history. For example, if the user has frequently used text input in the past, the system can suggest an input method optimized for text input. Similarly, if the user has frequently used voice input in the past, the system can suggest an input method optimized for voice input. Furthermore, if the user has frequently used image input in the past, the system can suggest an input method optimized for image input. This allows the system to suggest the optimal input method by referring to past input history. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that receives the user's past input data as input and suggests the optimal input method based on that data.
[0058] The following briefly describes the processing flow for example form 1.
[0059] Step 1: The photography team takes pictures of the educational building blocks. The photography team can take pictures of the building blocks using a smartphone or digital camera. They can also automatically optimize the arrangement and angle of the blocks when taking pictures, and can automatically identify the color and shape of the blocks and select the optimal shooting settings. Step 2: The analysis unit uses a generation AI to analyze the photographs taken by the photography unit. The analysis unit can analyze the types and arrangement of blocks using an image recognition algorithm. It can also analyze photographs taken with multiple different blocks mixed together and generate novel assembly methods using all of them. Furthermore, it can optimize the analysis results by taking into account the material and texture of the blocks. Step 3: The generation unit uses generation AI to generate original assembly instructions based on the results analyzed by the analysis unit. The generation unit can generate assembly instructions tailored to what the user wants to build. It can also suggest the optimal assembly method by referring to the user's past assembly history. Furthermore, it can customize the assembly instructions based on the user's current projects and areas of interest. Step 4: The input section allows the user to input what they want to create. The input section can accept text input or voice input to convey the user's wishes. It can also estimate the user's emotions and adjust the display of the input interface based on those emotions. Furthermore, it can suggest the optimal input method by referring to the user's past input history.
[0060] (Example of form 2) The educational block assembly support system according to an embodiment of the present invention is a system in which a generating AI generates original block assembly instructions for educational blocks taken from a photograph or a photograph taken on the spot. The educational block assembly support system works by having the user take a photograph of the educational blocks, and the generating AI analyzes the photograph to generate original assembly instructions. These instructions are generated to suit what the user wants to build (for example, a dinosaur or a castle). Furthermore, even if multiple different blocks are mixed and photographed, the system can generate a novel assembly method that incorporates all of them. This allows the user to gain new ideas and reuse the blocks. For example, the educational block assembly support system works by having the user take a photograph of the educational blocks. Next, the generating AI analyzes the photograph and generates original assembly instructions. These instructions are generated to suit what the user wants to build (for example, a dinosaur or a castle). Furthermore, even if multiple different blocks are mixed and photographed, the system can generate a novel assembly method that incorporates all of them. This allows the user to gain new ideas and reuse the blocks. Thus, the educational block assembly support system allows the user to gain new ideas and reuse the blocks.
[0061] The educational block assembly support system according to this embodiment comprises a shooting unit, an analysis unit, a generation unit, and an input unit. The shooting unit takes photographs of the educational blocks. The shooting unit can take photographs of the educational blocks using, for example, a smartphone or a digital camera. The shooting unit can also automatically optimize the arrangement and angle of the blocks when taking photographs. Furthermore, the shooting unit can automatically identify the color and shape of the blocks and select the optimal shooting settings. The analysis unit analyzes the photographs taken by the shooting unit using a generation AI. The analysis unit can analyze the types and arrangement of the blocks using, for example, an image recognition algorithm. Furthermore, the analysis unit can analyze photographs taken with multiple different blocks mixed together and generate novel assembly methods that combine them all. Furthermore, the analysis unit can optimize the analysis results by considering, for example, the material and texture of the blocks. The generation unit uses a generation AI to generate original assembly instructions based on the results analyzed by the analysis unit. The generation unit can generate assembly instructions tailored to what the user wants to make. Furthermore, the generation unit can suggest the optimal assembly method by referring to the user's past assembly history. Furthermore, the generation unit can customize the assembly instructions based on the user's current projects or areas of interest, for example. The input unit allows the user to input what they want to create. The input unit can input the user's wishes using, for example, text input or voice input. The input unit can also estimate the user's emotions and adjust the display method of the input interface based on the estimated emotions. Furthermore, the input unit can suggest the optimal input method by referring to the user's past input history, for example. As a result, the educational block assembly support system according to the embodiment allows the user to gain new ideas and reuse the blocks.
[0062] The photography unit takes pictures of the educational building blocks. The photography unit can take pictures of the blocks using, for example, a smartphone or digital camera. Specifically, it launches the camera app on a smartphone and takes pictures while checking the arrangement of the blocks. For digital cameras, it is recommended to use a tripod for stable shooting. The photography unit can also automatically optimize the arrangement and angle of the blocks before shooting. This utilizes the camera's autofocus function and AI-based image correction technology. Furthermore, the photography unit can automatically identify the color and shape of the blocks and select the optimal shooting settings. This allows the photography unit to accurately capture the detailed characteristics of the blocks and provide high-quality image data to the analysis unit. For example, it can automatically adjust the camera's white balance and exposure settings to ensure that the block colors are accurately reproduced. Also, even when multiple blocks are mixed together, the photography unit optimizes the position and angle of each block during shooting, enabling the analysis unit to process the data efficiently. This allows the photography unit to support users in easily taking high-quality photos and improve the overall accuracy and efficiency of the system.
[0063] The analysis unit uses generative AI to analyze photographs taken by the photography unit. For example, the analysis unit can analyze the type and arrangement of blocks using image recognition algorithms. Specifically, it uses a deep learning-based image recognition model to identify the shape, color, and size of the blocks. For instance, it can use a convolutional neural network (CNN) to analyze the edges and textures of the blocks and extract the characteristics of each block. The analysis unit can also analyze photographs of mixed blocks and generate novel assembly methods using all of them. This involves a process where the generative AI learns the block arrangement patterns and proposes new assembly methods. Furthermore, the analysis unit can optimize the analysis results by considering the material and texture of the blocks. For example, it can identify whether a block is made of plastic or wood and propose an assembly method accordingly. This allows the analysis unit to provide the optimal assembly method based on the type and characteristics of the blocks the user possesses. Additionally, the analysis unit can learn user preferences and tendencies based on past analysis data to provide more personalized analysis results. This enables the analysis unit to perform flexible analysis tailored to user needs and improve the overall user experience of the system.
[0064] The generation unit uses a generation AI to generate original assembly instructions based on the results analyzed by the analysis unit. For example, the generation unit can generate assembly instructions tailored to what the user wants to create. Specifically, when the user inputs the type or theme of their desired creation through the input unit, the generation AI generates the optimal assembly procedure based on that information. For example, if the user inputs that they want to build a "robot," the generation unit will generate a detailed instruction manual for assembling the robot based on the block types and placement information obtained from the analysis unit. The generation unit can also suggest the optimal assembly method by referring to the user's past assembly history. This includes saving data on previously created creations and customizing new assembly procedures based on that data. Furthermore, the generation unit can customize the assembly instructions based on the user's current projects or areas of interest. For example, if the user wants to create a "space" themed creation, the generation unit will generate instructions incorporating space-related block placements and designs. In this way, the generation unit provides assembly instructions tailored to the user's individual needs and interests, supporting them in enjoying the block building process more. Additionally, the generation unit can save the generated instructions in PDF or image format, allowing the user to access them at any time. This allows the generation unit to provide users with high-quality, customized assembly instructions, improving the overall convenience and satisfaction of the system.
[0065] The input section is where the user enters what they want to create. The input section allows users to input their wishes using methods such as text input or voice input. Specifically, users can input text using a smartphone or tablet keyboard, or communicate their wishes by voice using voice recognition technology. For example, if a user voice-inputs "I want to make a dinosaur," the voice recognition system converts that into text and transmits it to the system. Furthermore, the input section can estimate the user's emotions and adjust the display of the input interface based on those emotions. This includes a process of estimating emotions from the user's facial expressions and tone of voice using facial recognition and voice analysis technologies. For example, if the user appears happy, the interface's colors and design will be brightened; conversely, if they appear tired, it will be changed to a simpler, more calming design. Additionally, the input section can suggest the optimal input method by referring to the user's past input history. This includes a function that predicts frequently used phrases and keywords based on data the user has previously entered, assisting with input. For example, if a user has previously entered "I want to make a car," the keyword "car" will be automatically suggested the next time they enter a similar wish. This allows the input section to support users in easily and quickly entering their preferences, improving the overall usability of the system.
[0066] The generation unit can generate assembly instructions tailored to what the user wants to build. For example, if the user wants to build a dinosaur, the generation unit will generate assembly instructions for a dinosaur. It can also generate assembly instructions for a castle if the user wants to build one, and for example, if the user wants to build a car. This improves user satisfaction by generating assembly instructions that match the user's wishes. Some or all of the above processing in the generation unit is performed using a generative AI. For example, the generation unit can use a generative AI model that takes what the user wants to build as input and generates assembly instructions based on that input.
[0067] The analysis unit includes a mixed analysis unit that analyzes photographs taken with multiple different blocks mixed together and generates novel assembly methods using all of them. For example, the analysis unit can analyze photographs taken with blocks from different manufacturers mixed together. It can also analyze photographs taken with blocks of different shapes mixed together. Furthermore, it can analyze photographs taken with blocks of different colors mixed together. This allows for the generation of novel assembly methods using multiple different blocks, thereby providing new ideas. Some or all of the above processing in the analysis unit is performed using generative AI. For example, the analysis unit can use a generative AI model that takes a photograph taken with multiple different blocks mixed together as input and generates a novel assembly method based on that photograph.
[0068] The camera unit can estimate the user's emotions and adjust the shooting timing based on the estimated emotions. For example, if the user is excited, the AI can automatically speed up the shooting timing to ensure the moment isn't missed. Alternatively, if the user is relaxed, the AI can shoot at a slower pace to capture natural expressions. Furthermore, if the user is focused, the AI can time the shoot to capture the moment the blocks are placed perfectly. This allows for capturing the optimal moment by adjusting the shooting timing according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the camera unit is performed using AI. For example, the camera unit can use an AI model that receives user facial expression data as input and estimates emotions based on that data.
[0069] The camera unit can automatically optimize the placement and angle of blocks during shooting. For example, the camera unit can use AI to analyze the block placement in real time and shoot at the optimal angle. The camera unit can also use AI to automatically adjust the block placement to capture the most aesthetically pleasing composition. Furthermore, the camera unit can use AI to consider the position of the block shadows and shoot with optimal lighting. This allows for beautiful compositions by optimizing the block placement and angle. Some or all of the above processes in the camera unit are performed using AI. For example, the camera unit can use an AI model that receives block placement data as input and calculates the optimal angle based on that data.
[0070] The shooting unit can automatically identify the color and shape of the blocks during shooting and select the optimal shooting settings. For example, the shooting unit can use AI to identify the color of the blocks, set the optimal white balance, and take a picture. It can also use AI to identify the shape of the blocks, select the optimal focus settings, and take a picture. Furthermore, it can use AI to simultaneously identify the color and shape of the blocks, select the optimal exposure settings, and take a picture. This allows for the selection of optimal shooting settings according to the color and shape of the blocks, resulting in better photographs. Some or all of the above processing in the shooting unit is performed using AI. For example, the shooting unit can use an AI model that receives block color and shape data as input and calculates the optimal shooting settings based on that data.
[0071] The camera unit can estimate the user's emotions and determine the priority of which blocks to photograph based on the estimated emotions. For example, if the user is excited, the AI will prioritize photographing the blocks that the user is most interested in. Alternatively, if the user is relaxed, the AI can select blocks to photograph considering the overall balance. Furthermore, if the user is focused, the AI can prioritize photographing the blocks that the user is most focused on. This allows for the capture of photos that capture the user's interest by prioritizing blocks according to their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the camera unit is performed using AI. For example, the camera unit can use an AI model that receives user facial expression data as input and estimates emotions based on that data.
[0072] The camera unit can suggest the optimal shooting method by referring to the user's past shooting history during shooting. For example, the camera unit's AI can analyze the user's past shooting history and suggest the most successful shooting method. The camera unit can also suggest the optimal shooting angle based on the user's past shooting history using AI. Furthermore, the camera unit can automatically select the optimal shooting settings by referring to the user's past shooting history using AI. In this way, the camera unit can suggest the optimal shooting method by referring to past shooting history. Some or all of the above processes in the camera unit are performed using AI. For example, the camera unit can use an AI model that receives the user's past shooting data as input and suggests the optimal shooting method based on that data.
[0073] The shooting unit can select the optimal shooting settings by considering the user's surrounding environment during shooting. For example, the shooting unit can use AI to measure the ambient light intensity and select the optimal exposure setting. It can also use AI to measure the ambient sound level and select the optimal shutter speed. Furthermore, it can use AI to measure the ambient temperature and select the optimal white balance. This allows for the selection of optimal shooting settings by considering the surrounding environment. Some or all of the above processing in the shooting unit is performed using AI. For example, the shooting unit can use an AI model that receives ambient environmental data as input and calculates the optimal shooting settings based on that data.
[0074] The analysis unit can estimate the user's emotions and adjust the accuracy of the analysis based on the estimated emotions. For example, if the user is excited, the generative AI can improve the accuracy of the analysis and provide detailed results. Similarly, if the user is relaxed, the generative AI can adjust the accuracy of the analysis and provide results that emphasize overall balance. Furthermore, if the user is focused, the generative AI can optimize the accuracy of the analysis and provide results that focus on specific aspects. This allows for more appropriate analysis results by adjusting the accuracy of the analysis according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the analysis unit are performed using generative AI. For example, the analysis unit can use a generative AI model that receives user emotion data as input and adjusts the accuracy of the analysis based on that data.
[0075] The analysis unit can optimize the analysis results by considering the material and texture of the block during the analysis. For example, the analysis unit's generating AI can identify the material of the block and select the optimal analysis algorithm. The analysis unit can also, for example, have the generating AI analyze the texture of the block and set the optimal analysis parameters. Furthermore, the analysis unit can, for example, have the generating AI consider both the material and texture of the block simultaneously to provide the optimal analysis results. This allows for more accurate analysis results by considering the material and texture of the block. Some or all of the above processes in the analysis unit are performed using the generating AI. For example, the analysis unit can use a generating AI model that receives block material and texture data as input and provides the optimal analysis results based on that data.
[0076] The analysis unit can automatically generate block combination patterns during analysis and propose the optimal combination. For example, the analysis unit's generating AI can generate the optimal combination pattern based on the shape and color of the blocks. The analysis unit can also propose the optimal combination pattern by considering the arrangement of the blocks, for example. Furthermore, the analysis unit can provide the optimal combination pattern by analyzing the frequency of block usage, for example. In this way, the analysis unit can propose the optimal combination by automatically generating block combination patterns. Some or all of the above processes in the analysis unit are performed using the generating AI. For example, the analysis unit can use a generating AI model that receives block shape and color data as input and generates the optimal combination pattern based on that data.
[0077] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated user emotions. For example, if the user is excited, the generation AI can provide a visually stimulating display method. The analysis unit can also provide a calming display method if the user is relaxed. Furthermore, if the user is focused, the generation AI can provide a display method that includes detailed information. By adjusting the display method according to the user's emotions, more appropriate analysis results can be provided. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit is performed using generation AI. For example, the analysis unit can use a generation AI model that receives user emotion data as input and adjusts the display method based on that data.
[0078] The analysis unit can select the optimal analysis method by referring to the user's past assembly history during analysis. For example, the analysis unit's generating AI can analyze the user's past assembly history and select the optimal analysis algorithm. The analysis unit can also, for example, have the generating AI set the optimal analysis parameters based on the user's past assembly history. Furthermore, the analysis unit can, for example, have the generating AI refer to the user's past assembly history and provide the optimal analysis results. In this way, the optimal analysis method can be selected by referring to the past assembly history. Some or all of the above processes in the analysis unit are performed using the generating AI. For example, the analysis unit can use a generating AI model that receives the user's past assembly data as input and selects the optimal analysis method based on that data.
[0079] The analysis unit can customize the analysis results based on the user's current projects and areas of interest during the analysis process. For example, the analysis unit can use a generative AI to analyze the user's current projects and provide optimal analysis results. Alternatively, the analysis unit can use a generative AI to consider the user's areas of interest and provide customized analysis results. Furthermore, the analysis unit can use a generative AI to provide optimal analysis results based on the user's project progress. This allows for the provision of more appropriate analysis results by customizing them based on the current projects and areas of interest. Some or all of the above-described processes in the analysis unit are performed using generative AI. For example, the analysis unit can use a generative AI model that receives the user's project data as input and customizes the analysis results based on that data.
[0080] The generation unit can estimate the user's emotions and adjust the presentation of the generated assembly instructions based on the estimated user emotions. For example, if the user is excited, the generation AI can provide a visually stimulating presentation. Similarly, if the user is relaxed, the generation AI can provide a calm presentation. Furthermore, if the user is focused, the generation AI can provide a presentation containing detailed information. This allows for the provision of more appropriate assembly instructions by adjusting the presentation according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generation AI. The generation AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the generation unit are performed using generation AI. For example, the generation unit can use a generation AI model that receives user emotion data as input and adjusts the presentation based on that data.
[0081] The generation unit can propose the optimal assembly method by referring to the user's past assembly history during generation. For example, the generation unit's generation AI can analyze the user's past assembly history and propose the optimal assembly method. The generation unit can also provide the optimal assembly procedure based on the user's past assembly history, for example. Furthermore, the generation unit can propose the optimal assembly pattern by referring to the user's past assembly history, for example. In this way, the optimal assembly method can be proposed by referring to past assembly history. Some or all of the above processing in the generation unit is performed using the generation AI. For example, the generation unit can use a generation AI model that receives the user's past assembly data as input and proposes the optimal assembly method based on that data.
[0082] The generation unit can customize the assembly instructions based on the user's current projects and areas of interest during the generation process. For example, the generation unit's generation AI can analyze the user's current projects and provide customized assembly instructions. Alternatively, the generation unit can use its generation AI to consider the user's areas of interest and provide optimal assembly instructions. Furthermore, the generation unit can use its generation AI to provide customized assembly instructions based on the user's project progress. This allows for the provision of more appropriate assembly instructions by customizing them based on the current projects and areas of interest. Some or all of the above-described processes in the generation unit are performed using generation AI. For example, the generation unit can use a generation AI model that receives user project data as input and customizes the assembly instructions based on that data.
[0083] The generation unit can estimate the user's emotions and determine the priority of the assembly instructions to generate based on the estimated user emotions. For example, if the user is excited, the generation AI will prioritize generating the assembly instructions that are most interesting to the user. Also, if the user is relaxed, the generation AI can generate assembly instructions considering the overall balance. Furthermore, if the user is focused, the generation AI can prioritize generating the most detailed assembly instructions. In this way, by determining priorities according to the user's emotions, it is possible to provide assembly instructions that are interesting to the user. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generation AI. The generation AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the generation unit is performed using a generation AI. For example, the generation unit can use a generation AI model that takes user emotion data as input and determines priorities based on that data.
[0084] The generation unit can propose an optimal assembly method while considering the user's geographical location information during generation. For example, the generation unit's generation AI can propose an optimal assembly method based on the user's geographical location information. Furthermore, the generation unit can also provide region-specific assembly methods by considering the user's geographical location information. Additionally, the generation unit can propose an optimal assembly procedure by referencing the user's geographical location information. This allows for the provision of region-specific assembly methods by considering geographical location information. Some or all of the above-described processes in the generation unit are performed using the generation AI. For example, the generation unit can use a generation AI model that receives the user's geographical location data as input and proposes an optimal assembly method based on that data.
[0085] The generation unit can analyze the user's social media activity during generation and propose relevant assembly methods. For example, the generation unit's generation AI can analyze the user's social media activity and propose the optimal assembly method. The generation unit can also provide relevant assembly steps based on the user's social media activity, for example. Furthermore, the generation unit can refer to the user's social media activity and propose the optimal assembly pattern, for example. This allows for the provision of assembly methods based on the user's interests by analyzing social media activity. Some or all of the above processing in the generation unit is performed using a generation AI. For example, the generation unit can use a generation AI model that receives the user's social media data as input and proposes the optimal assembly method based on that data.
[0086] The input unit can estimate the user's emotions and adjust the display method of the input interface based on the estimated emotions. For example, if the user is excited, the generative AI can provide a visually stimulating display method. Similarly, if the user is relaxed, the generative AI can provide a calming display method. Furthermore, if the user is focused, the generative AI can provide a display method containing detailed information. This allows for a more appropriate input interface by adjusting the display method according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the input unit is performed using generative AI. For example, the input unit can use a generative AI model that receives user emotion data as input and adjusts the display method based on that data.
[0087] The input unit can suggest the optimal input method by referring to the user's past input history during input. For example, the input unit's generative AI can analyze the user's past input history and suggest the optimal input method. The input unit can also provide the optimal input procedure based on the user's past input history, for example. Furthermore, the input unit can suggest the optimal input pattern by referring to the user's past input history, for example. In this way, the optimal input method can be suggested by referring to past input history. Some or all of the above processing in the input unit is performed using generative AI. For example, the input unit can use a generative AI model that receives the user's past input data as input and suggests the optimal input method based on that data.
[0088] The input unit can customize the input content based on the user's current projects and areas of interest during input. For example, the input unit can use a generative AI to analyze the user's current projects and provide customized input content. Alternatively, the input unit can use a generative AI to consider the user's areas of interest and provide optimal input content. Furthermore, the input unit can use a generative AI to provide customized input content based on the user's project progress. This allows for the provision of more appropriate input content by customizing it based on the current projects and areas of interest. Some or all of the above processing in the input unit is performed using generative AI. For example, the input unit can use a generative AI model that receives the user's project data as input and customizes the input content based on that data.
[0089] The input unit can estimate the user's emotions and prioritize input content based on the estimated emotions. For example, if the user is excited, the generative AI will prioritize displaying the input content that is most interesting to the user. Similarly, if the user is relaxed, the generative AI can display input content considering the overall balance. Furthermore, if the user is focused, the generative AI can prioritize displaying the input content that is most detailed. This allows for the provision of user-interesting input content by prioritizing content according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the input unit is performed using generative AI. For example, the input unit can use a generative AI model that receives user emotion data as input and determines priorities based on that data.
[0090] The input unit can propose the optimal input method while considering the user's device information. For example, the input unit's generative AI can propose the optimal input method based on the user's device information. Furthermore, the input unit can, for example, have the generative AI consider the user's device information and provide an input procedure optimized for the device. Additionally, the input unit can, for example, have the generative AI refer to the user's device information and propose the optimal input pattern. This allows for the proposal of the optimal input method by considering the device information. Some or all of the above processing in the input unit is performed using generative AI. For example, the input unit can use a generative AI model that receives the user's device information as input and proposes the optimal input method based on that data.
[0091] The input unit can analyze the user's social media activity during input and suggest relevant input content. For example, the input unit can use a generative AI to analyze the user's social media activity and suggest the optimal input content. The input unit can also use a generative AI to provide relevant input steps based on the user's social media activity. Furthermore, the input unit can use a generative AI to refer to the user's social media activity and suggest the optimal input pattern. This allows for the provision of input content based on the user's interests by analyzing social media activity. Some or all of the above processing in the input unit is performed using a generative AI. For example, the input unit can use a generative AI model that receives the user's social media data as input and suggests the optimal input content based on that data.
[0092] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0093] The educational building block assembly support system can estimate the user's emotions and adjust the difficulty level of the assembly instructions based on those emotions. For example, if the user is excited, the system will provide a more difficult assembly instruction to stimulate their desire to challenge themselves. If the user is relaxed, the system will provide an easier assembly instruction to allow them to enjoy the activity in a relaxed state. Furthermore, if the user is focused, the system will provide an assembly instruction of moderate difficulty to allow them to enjoy the activity while maintaining their concentration. This enables difficulty level adjustment according to the user's emotions, providing a more appropriate building experience. Emotion estimation is achieved using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the system is performed using generative AI. For example, the system can use an AI model that receives user facial expression data as input and estimates emotions based on that data.
[0094] The educational building block assembly support system can refer to the user's past assembly history and suggest new assembly methods based on the user's previously created works. For example, if the user previously created a dinosaur, the system can suggest a new variation of that dinosaur. Also, if the user previously created a castle, the system can suggest a new castle design based on that castle. Furthermore, if the user previously created a car, the system can suggest a new car model based on that car. In this way, the system can leverage the user's past assembly history to provide new ideas. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that takes the user's past assembly data as input and suggests new assembly methods based on that data.
[0095] The educational building block assembly support system can estimate the user's emotions and adjust the display method of the assembly instructions based on those emotions. For example, if the user is excited, the system will provide a visually stimulating display method to maintain the user's excitement. If the user is relaxed, the system will provide a calm display method to allow them to enjoy building in a relaxed state. Furthermore, if the user is focused, the system will provide a display method that includes detailed information, allowing them to continue building while maintaining their concentration. In this way, by adjusting the display method according to the user's emotions, a more appropriate building experience can be provided. Emotion estimation is achieved using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the system is performed using generative AI. For example, the system can use an AI model that receives user facial expression data as input and adjusts the display method based on that data.
[0096] The educational building block assembly support system can suggest region-specific assembly methods by considering the user's geographical location. For example, if the user is in Japan, the system can suggest assembly methods based on traditional Japanese architecture and culture. If the user is in the United States, the system can suggest assembly methods based on historical American buildings and culture. Furthermore, if the user is in Europe, the system can suggest assembly methods based on European architecture such as castles and churches. In this way, by considering geographical location, the system can provide region-specific assembly methods. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that takes the user's geographical location data as input and suggests the optimal assembly method based on that data.
[0097] The educational building block assembly support system can estimate the user's emotions and prioritize the assembly instructions based on those emotions. For example, if the user is excited, the system will prioritize generating the most interesting assembly instructions. If the user is relaxed, the system can also generate instructions considering the overall balance. Furthermore, if the user is focused, the system can prioritize generating the most detailed assembly instructions. This allows the system to provide assembly instructions that are interesting to the user by prioritizing them according to their emotions. Emotion estimation is achieved using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that takes user emotion data as input and determines priorities based on that data.
[0098] The educational building block assembly support system can analyze a user's social media activity and suggest relevant building methods. For example, if a user frequently posts about dinosaurs on social media, the system can suggest how to build a dinosaur. Similarly, if a user frequently posts about castles, the system can suggest how to build a castle. Furthermore, if a user frequently posts about cars, the system can suggest how to build a car. In this way, by analyzing social media activity, the system can provide building methods based on the user's interests. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that takes the user's social media data as input and suggests the optimal building method based on that data.
[0099] The educational building block assembly support system can estimate the user's emotions and adjust the presentation of the assembly instructions based on those emotions. For example, if the user is excited, the system can provide a visually stimulating presentation. If the user is relaxed, the system can provide a calm presentation. Furthermore, if the user is focused, the system can provide a presentation that includes detailed information. By adjusting the presentation according to the user's emotions, the system can provide more appropriate assembly instructions. Emotion estimation is achieved using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that takes user emotion data as input and adjusts the presentation based on that data.
[0100] The educational building block assembly support system can suggest the optimal input method by considering the user's device information. For example, if the user is using a smartphone, the system can suggest an input method optimized for touchscreens. If the user is using a tablet, the system can also suggest an input method optimized for large screens. Furthermore, if the user is using a desktop computer, the system can suggest an input method optimized for keyboards and mice. In this way, the system can suggest the optimal input method by considering device information. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that receives the user's device information as input and suggests the optimal input method based on that data.
[0101] The educational building block assembly support system can estimate the user's emotions and prioritize the assembly instructions based on those emotions. For example, if the user is excited, the system will prioritize generating the most interesting assembly instructions. If the user is relaxed, the system can also generate instructions considering the overall balance. Furthermore, if the user is focused, the system can prioritize generating the most detailed assembly instructions. This allows the system to provide assembly instructions that are interesting to the user by prioritizing them according to their emotions. Emotion estimation is achieved using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that takes user emotion data as input and determines priorities based on that data.
[0102] The educational block assembly support system can suggest the optimal input method by referring to the user's past input history. For example, if the user has frequently used text input in the past, the system can suggest an input method optimized for text input. Similarly, if the user has frequently used voice input in the past, the system can suggest an input method optimized for voice input. Furthermore, if the user has frequently used image input in the past, the system can suggest an input method optimized for image input. This allows the system to suggest the optimal input method by referring to past input history. Some or all of the above processing in the system is performed using generative AI. For example, the system can use a generative AI model that receives the user's past input data as input and suggests the optimal input method based on that data.
[0103] The following briefly describes the processing flow for example form 2.
[0104] Step 1: The photography team takes pictures of the educational building blocks. The photography team can take pictures of the building blocks using a smartphone or digital camera. They can also automatically optimize the arrangement and angle of the blocks when taking pictures, and can automatically identify the color and shape of the blocks and select the optimal shooting settings. Step 2: The analysis unit uses a generation AI to analyze the photographs taken by the photography unit. The analysis unit can analyze the types and arrangement of blocks using an image recognition algorithm. It can also analyze photographs taken with multiple different blocks mixed together and generate novel assembly methods using all of them. Furthermore, it can optimize the analysis results by taking into account the material and texture of the blocks. Step 3: The generation unit uses generation AI to generate original assembly instructions based on the results analyzed by the analysis unit. The generation unit can generate assembly instructions tailored to what the user wants to build. It can also suggest the optimal assembly method by referring to the user's past assembly history. Furthermore, it can customize the assembly instructions based on the user's current projects and areas of interest. Step 4: The input section allows the user to input what they want to create. The input section can accept text input or voice input to convey the user's wishes. It can also estimate the user's emotions and adjust the display of the input interface based on those emotions. Furthermore, it can suggest the optimal input method by referring to the user's past input history.
[0105] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0106] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0107] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0108] Each of the multiple elements described above, including the shooting unit, analysis unit, generation unit, and input unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the shooting unit takes a photograph of the educational building blocks using the camera 42 of the smart device 14. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12 and analyzes the photograph. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12 and generates an original assembly instruction manual based on the analysis results. The input unit inputs the user's wishes using the receiving device 38 of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0109] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0110] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0111] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0112] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0113] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0114] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0115] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0116] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0117] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0118] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0119] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0120] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0121] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0122] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0123] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0124] Each of the multiple elements described above, including the imaging unit, analysis unit, generation unit, and input unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the imaging unit uses the camera 42 of the smart glasses 214 to take a picture of the educational building blocks. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12 to analyze the captured picture. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12 to generate an original assembly instruction manual based on the analysis results. The input unit uses the microphone 238 of the smart glasses 214 to input the user's wishes. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0125] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0126] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0127] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0128] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0129] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0130] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0131] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0132] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0133] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0134] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0135] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0136] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0137] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0138] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0139] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0140] Each of the multiple elements described above, including the shooting unit, analysis unit, generation unit, and input unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the shooting unit takes a photograph of the educational building blocks using the camera 42 of the headset terminal 314. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12 and analyzes the photograph. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12 and generates an original assembly instruction manual based on the analysis results. The input unit inputs the user's wishes using the microphone 238 of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0141] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0142] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0143] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0144] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0145] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0146] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0147] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0148] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0149] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0150] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0151] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0152] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0153] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0154] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0155] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0156] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0157] Each of the multiple elements described above, including the imaging unit, analysis unit, generation unit, and input unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the imaging unit takes pictures of the educational blocks using the camera 42 of the robot 414. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12 and analyzes the captured pictures. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12 and generates original assembly instructions based on the analysis results. The input unit inputs the user's wishes using the microphone 238 of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0158] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0159] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0160] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0161] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0162] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0163] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0164] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0165] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0166] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0167] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0168] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0169] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0170] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0171] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0172] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0173] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0174] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0175] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0176] (Note 1) The photography team takes pictures of educational building blocks, An analysis unit that analyzes the photographs taken by the aforementioned photography unit, A generation unit that generates an original assembly instruction manual based on the results of the analysis performed by the aforementioned analysis unit, It includes an input section where the user inputs what they want to create. A system characterized by the following features. (Note 2) The generating unit is Generates assembly instructions tailored to what the user wants to create. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned analysis unit, It features a mixed analysis unit that analyzes photographs taken with multiple different blocks mixed together and generates a novel assembly method that combines all of them. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned imaging unit is It estimates the user's emotions and adjusts the shooting timing based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned imaging unit is The block placement and angle are automatically optimized during shooting. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned imaging unit is During shooting, the system automatically identifies the color and shape of the blocks and selects the optimal shooting settings. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned imaging unit is It estimates the user's emotions and determines the priority of the blocks to photograph based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned imaging unit is During shooting, the system refers to the user's past shooting history to suggest the optimal shooting method. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned imaging unit is During shooting, the system selects the optimal shooting settings by considering the user's surrounding environment. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned analysis unit, It estimates the user's emotions and adjusts the accuracy of the analysis based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned analysis unit, During analysis, the analysis results are optimized by considering the material and texture of the blocks. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit, During analysis, the system automatically generates block combination patterns and proposes the optimal combination. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During analysis, the system selects the optimal analysis method by referring to the user's past assembly history. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, the results are customized based on the user's current projects and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 16) The generating unit is We estimate the user's emotions and adjust the way the assembly instructions are presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The generating unit is During generation, the system will refer to the user's past assembly history to suggest the optimal assembly method. The system described in Appendix 1, characterized by the features described herein. (Note 18) The generating unit is During generation, the assembly instructions are customized based on the user's current projects and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 19) The generating unit is It estimates the user's emotions and determines the priority of the assembly instructions to be generated based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The generating unit is During generation, the system proposes the optimal assembly method, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 21) The generating unit is During generation, the system analyzes the user's social media activity and suggests relevant assembly methods. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned input unit is It estimates the user's emotions and adjusts how the input interface is displayed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned input unit is During input, the system refers to the user's past input history to suggest the optimal input method. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned input unit is When users enter data, the input content is customized based on their current projects and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned input unit is It estimates the user's emotions and prioritizes input content based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned input unit is When inputting data, the system suggests the optimal input method, taking into account the user's device information. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned input unit is During input, the system analyzes the user's social media activity and suggests relevant input content. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0177] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. The photography team takes pictures of educational building blocks, An analysis unit that analyzes the photographs taken by the aforementioned photography unit, A generation unit that generates an original assembly instruction manual based on the results of the analysis performed by the aforementioned analysis unit, It includes an input section where the user inputs what they want to create. A system characterized by the following features.
2. The generating unit is Generates assembly instructions tailored to what the user wants to create. The system according to feature 1.
3. The aforementioned analysis unit, It features a mixed analysis unit that analyzes photographs taken with multiple different blocks mixed together and generates a novel assembly method that combines all of them. The system according to feature 1.
4. The aforementioned imaging unit is It estimates the user's emotions and adjusts the shooting timing based on the estimated user emotions. The system according to feature 1.
5. The aforementioned imaging unit is The block placement and angle are automatically optimized during shooting. The system according to feature 1.
6. The aforementioned imaging unit is During shooting, the system automatically identifies the color and shape of the blocks and selects the optimal shooting settings. The system according to feature 1.
7. The aforementioned imaging unit is It estimates the user's emotions and determines the priority of the blocks to photograph based on the estimated user emotions. The system according to feature 1.
8. The aforementioned imaging unit is During shooting, the system refers to the user's past shooting history to suggest the optimal shooting method. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A