System

A system simplifies VR game animation creation by allowing users to select and edit animation atmospheres and scenes, using a generative AI model to generate and save animations, addressing complexity and cost issues in VR game development.

JP2026024087APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126408
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Creating character animations in VR game development is an extremely complex, time-consuming, and costly task, requiring full-body tracking and motion capture, leading to schedule delays and a heavy burden on developers.

Method used

A system that allows users to select an animation atmosphere and usage scene, generates predetermined animation data, provides it for editing, and saves the edited data, utilizing a generative AI model to simplify the process.

Benefits of technology

Enables users to easily and efficiently create and edit high-quality character animations for VR games without specialized knowledge, reducing developer workload and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026024087000001_ABST
    Figure 2026024087000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for allowing a user to select an animation mood and a usage scene; means for generating predetermined animation data based on the selected mood and usage scene; means for providing the generated animation data to the user; means for allowing the user to edit the provided animation data; and means for storing the edited animation data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, creating character animations in VR game development has become an extremely complex, time-consuming, and costly task. Even basic movements like walking and running require full-body tracking, motion capture, and even modeling from scratch, consuming significant resources. This has led to schedule delays for VR game development projects and placed a heavy burden on developers. Given this background, there is a demand for tools that allow for easy and efficient animation creation. [Means for solving the problem]

[0005] The present invention solves the above problem by providing a system including: means for a user to select an animation atmosphere and usage scene; means for generating predetermined animation data based on the selected atmosphere and usage scene; means for providing the generated animation data to the user; means for the user to edit the provided animation data; and means for saving the edited animation data.

[0006] A "user" is an individual or group that uses the system and is the entity that generates, checks, edits, and saves animation data.

[0007] "Animation" is a series of images or object changes that represent the movement of a character, and is used as part of a game or video content.

[0008] "Atmosphere" is an attribute that defines the emotion or style of animation data, and includes "formal," "casual," "playful," and the like.

[0009] "Usage scene" refers to a specific situation or context in which the animation data is used, and includes "battle," "fantasy," "everyday life," etc.

[0010] "Means" refers to a method or device by which a system achieves a specific function.

[0011] "Generation" refers to the process of creating new animation data based on input conditions.

[0012] "Providing" refers to the act of displaying or transmitting the generated animation data to the user.

[0013] "Editing" refers to the act of a user making changes or adjustments to the animation data provided.

[0014] "Saving" refers to the act of recording animation data edited by the user in the system so that it can be reused or re-edited later.

[0015] The term "system" refers to the overall configuration of hardware and software necessary to implement the present invention. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention is a system for easily and efficiently creating character animations for VR games, and is implemented as follows.

[0038] Receiving User Input

[0039] Terminal

[0040] When a user launches the application, an initial screen is displayed, which provides an interface for selecting the animation atmosphere and usage scenario.

[0041] The user selects the preferred atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy).

[0042] The selected information is packaged in JSON format and sent to the server.

[0043] Generate animation

[0044] server

[0045] The input data received from the terminal is analyzed, and appropriate animation data is selected based on the selected atmosphere and usage scene.

[0046] The server searches the preset motions in the database and selects a motion that meets the conditions.

[0047] The selected motions are combined to generate a series of animation sequences.

[0048] The generated animation sequence is sent to the device in JSON format.

[0049] Sending Animations

[0050] server

[0051] The generated animation data is sent to the device as an HTTP response and is reflected to the user in real time.

[0052] The server will continue to maintain the connection after sending to allow for editing and resending.

[0053] User viewing and editing of animations

[0054] Terminal

[0055] The user can check the animation movement on the screen based on the received animation data.

[0056] You can use the editing tools provided to fine-tune motion joins and timelines.

[0057] The edited animation data is sent back to the server in JSON format.

[0058] Saving and resubmitting animations

[0059] server

[0060] Upon receiving the edited animation data, the server stores the data in a database.

[0061] If necessary, the saved data is sent again to the terminal so that the user can make a final confirmation.

[0062] Specific examples

[0063] 1. Receiving User Input

[0064] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[0065] 2. Device: Sends selection to server.

[0066] 2. Creating and sending animations

[0067] 3. Server: Searches and generates casual, everyday walking motions based on the received data.

[0068] 4. Server: Sends the generated animation to the device.

[0069] 3. Check and edit the animation

[0070] 5. Terminal: Display the animation sent.

[0071] 6. User: Fine-tune any parts of the motion where the joints are not smooth.

[0072] 7. Terminal: Re-send the edited animation data.

[0073] 4. Save and resubmit animations

[0074] 8. Server: Receives the edited data and stores it in the database.

[0075] 9. Server: Re-sends the saved data to the device for final confirmation.

[0076] This allows users to easily and efficiently generate and edit character animations for VR games, helping to reduce developer workload and costs.

[0077] The processing flow will be explained below.

[0078] Step 1:

[0079] User: Launches the application and an interface appears on the screen for selecting the atmosphere and usage scenario.

[0080] Step 2:

[0081] User: Select "Casual" as the animation atmosphere and "Everyday" as the usage scene.

[0082] Step 3:

[0083] Terminal: Converts the selected information into JSON format and creates an HTTP request to send to the server.

[0084] Step 4:

[0085] On the device: Send the selected atmosphere and usage scene information to the server using an HTTP POST request.

[0086] Step 5:

[0087] Server: Analyzes the request received from the device and extracts the selected mood and usage scenario from the JSON data.

[0088] Step 6:

[0089] Server: Accesses the database and searches for appropriate preset motions based on the extracted atmosphere and usage scene.

[0090] Step 7:

[0091] Server: Selects appropriate motions from the search results and combines multiple motions to generate smooth animation sequences.

[0092] Step 8:

[0093] Server: Converts the generated animation sequence into JSON format and creates an HTTP response to send to the device.

[0094] Step 9:

[0095] Server: Sends the generated animation data to the device as an HTTP response.

[0096] Step 10:

[0097] Terminal: Receives the response from the server, parses the JSON data, and retrieves the animation.

[0098] Step 11:

[0099] Terminal: Displays animation on the screen based on the analyzed animation data.

[0100] Step 12:

[0101] User: Check the displayed animation and fine-tune the motion joints and timeline as necessary.

[0102] Step 13:

[0103] Terminal: Generates updated animation data reflecting the edits made by the user.

[0104] Step 14:

[0105] On the device: Convert the updated animation data into JSON format and create an HTTP request to send it back to the server.

[0106] Step 15:

[0107] On the device: Send the edited animation data to the server using an HTTP POST request.

[0108] Step 16:

[0109] Server: Receives the updated data sent from the device and parses the JSON data again.

[0110] Step 17:

[0111] Server: Saves the parsed update data to the database.

[0112] Step 18:

[0113] Server: Converts the saved update data into JSON format and creates an HTTP response to resend to the device.

[0114] Step 19:

[0115] Server: Re-sends the saved data to the device for final confirmation.

[0116] Step 20:

[0117] Terminal: Receives the final confirmation data from the server and displays it on the user interface.

[0118] Step 21:

[0119] User: Perform a final check of the animation and make further adjustments if necessary.

[0120] Example 1

[0121] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0122] Conventional character animation creation systems for VR games often require users to perform complex operations and require specialized knowledge, making it difficult to create animations efficiently and easily. Another issue is the inconsistency in the quality of the generated animations.

[0123] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0124] In this invention, the server includes means for analyzing input data using a generative AI model and generating an optimal animation sequence, means for transmitting and receiving JSON format data between the terminal and the server, and means for the user to select the atmosphere and usage scene of the animation. This allows users to efficiently and easily create and edit high-quality animations without requiring specialized knowledge, and also makes it possible to easily save and reuse the data.

[0125] "User" means an individual or organization that uses the system to create and edit character animations for VR games.

[0126] "Animation atmosphere" refers to the style and tone of character animation, and can include casual, formal, playful, etc.

[0127] "Usage scene" refers to an element that indicates the situation or background in which the character animation unfolds, and specifically includes everyday life, battle, fantasy, etc.

[0128] The term "means" refers to a specific mechanism or technology that realizes the functions or methods necessary to carry out the present invention.

[0129] "Generative AI models" refer to artificial intelligence algorithms and programs that automatically generate optimal animation sequences based on user input.

[0130] The "JSON format" is a data format used when sending and receiving data between a terminal and a server, and has a lightweight structure that is easy for humans to read.

[0131] "Terminal" refers to a device that a user uses to access the system, and refers to the hardware on which applications are executed.

[0132] "Server" refers to the central processing unit that analyzes the data sent by the user and generates and manages the animation sequences.

[0133] "Preset motion" refers to predefined standard character movements and actions, and is the basic element for generating animation.

[0134] An "animation sequence" refers to a series of character actions created by combining multiple motions.

[0135] "Database" refers to a storage system for saving generated and edited animation data and reusing it as needed.

[0136] "Editing tools" refers to software features that allow a user to adjust and modify the details of an animation.

[0137] The above definitions clearly define the terms that are included in the claims.

[0138] This invention is a system that allows users to easily and efficiently create and edit character animations for VR games. This system features a series of processes for generating animations based on the atmosphere and usage scene selected by the user.

[0139] Receiving User Input

[0140] Device:

[0141] When a user launches the application, an initial screen is displayed, which provides an interface for selecting the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy). Once the user selects the desired option, the information is packaged in JSON format and sent to the server.

[0142] Sending an animation generation request

[0143] server:

[0144] The system receives and analyzes the JSON data sent from the device. A generative AI model is used to analyze this data. Specifically, it searches a database for appropriate preset motions based on the user's selection, and then selects motions that match the conditions.

[0145] Generate animation

[0146] server:

[0147] Multiple preset motions in the database are combined to generate a series of animation sequences, taking into consideration the smooth connection of the selected motions. The generated animation sequence is then converted back to JSON format and prepared for transmission to the device.

[0148] Sending Animations

[0149] server:

[0150] The generated animation data is sent to the terminal as an HTTP response. This transmission is performed in real time, and the connection is maintained afterwards to allow editing and retransmission.

[0151] User viewing and editing of animations

[0152] Device:

[0153] The received animation data is displayed on the screen, allowing the user to check the movement. Furthermore, the user can fine-tune the motion joints and timeline using the provided editing tools. The edited animation data is then sent back to the server in JSON format.

[0154] Saving and resubmitting animations

[0155] server:

[0156] Once the edited animation data is received, it is saved in the database. If necessary, the saved animation data can be sent back to the terminal so that the user can make a final check.

[0157] This allows users to efficiently and easily generate and edit high-quality character animations without requiring specialized knowledge. Furthermore, the system makes it easy to save and reuse data, thereby contributing to reducing developer workload and costs.

[0158] Specific examples

[0159] Receiving User Input

[0160] 1. User: Launch the app and select the atmosphere "Casual" and the usage scene "Everyday."

[0161] 2. Terminal: Sends the selection to the server in JSON format.

[0162] Generate and send animations

[0163] 3. Server: Searches and generates casual, everyday walking motions based on the received data.

[0164] 4. Server: Sends the generated animation to the device in JSON format.

[0165] Checking and editing animations

[0166] 5. Terminal: Display the animation sent.

[0167] 6. User: Fine-tune any parts of the motion where the joins are not smooth.

[0168] 7. Device: The edited animation data is sent back to the server in JSON format.

[0169] Saving and resubmitting animations

[0170] 8. Server: Receives the edited data and stores it in the database.

[0171] 9. Server: Retransmits the saved data to the device for final confirmation.

[0172] Prompt Sentence Examples

[0173] "Please generate character animations for a VR game with a 'casual' atmosphere and an 'everyday' usage scenario. Please focus on walking motions."

[0174] As a result, this system provides users with an intuitive and easy-to-use interface and advanced animation generation functions.

[0175] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0176] Step 1:

[0177] User: When the application is launched, an initial screen appears, where the user selects the animation mood (e.g., casual, formal, playful) and usage scenario (e.g., everyday life, combat, fantasy).

[0178] Input: User-selected mood and usage scene.

[0179] Specific operation: The user selects "casual" and "everyday."

[0180] Output: The selected information is prepared in the terminal in JSON format.

[0181] Step 2:

[0182] Terminal: Sends the selected information to the server.

[0183] Input: Atmosphere and usage scene data packaged in JSON format.

[0184] Specific operation: The device sends JSON data to the server as an HTTP POST request.

[0185] Output: The server receives the JSON data.

[0186] Step 3:

[0187] Server: Parses the received JSON data and uses a generative AI model to set conditions to generate an appropriate animation sequence from the input data.

[0188] Input: JSON data received from the terminal.

[0189] Specific operation: The server parses the JSON data and searches the database for preset motions that correspond to "casual" and "everyday."

[0190] Output: Get multiple preset motions as search results.

[0191] Step 4:

[0192] Server: Selects preset motions that match the search criteria and combines them to generate a smooth animation sequence.

[0193] Input: Multiple preset motions retrieved from the database.

[0194] Specific operation: The server combines selected preset motions to generate a "casual everyday walking scene."

[0195] Output: The generated animation sequence is structured in JSON format.

[0196] Step 5:

[0197] Server: Sends the generated animation data to the terminal.

[0198] Input: JSON data of the generated animation sequence.

[0199] Specific operation: The server sends the animation JSON data to the device as an HTTP response.

[0200] Output: The device receives the animation sequence.

[0201] Step 6:

[0202] Terminal: The received animation data is displayed on the screen so that the user can check the movement.

[0203] Input: JSON data of the animation sequence received from the server.

[0204] Specific operation: The device plays the animation and provides an interface for the user to confirm.

[0205] Output: The user sees the animation visually.

[0206] Step 7:

[0207] User: Use the provided editing tools to fine-tune the animation, adjusting motion joins and timelines.

[0208] Input: Animation displayed on the device.

[0209] Specific operation: The user adjusts the animation joints by dragging and dropping and checks the smoothness.

[0210] Output: The edited animation is prepared in JSON format on the device.

[0211] Step 8:

[0212] Terminal: Send the edited animation data back to the server.

[0213] Input: JSON data of the user-edited animation sequence.

[0214] Specific operation: The device sends the edited JSON data to the server as an HTTP POST request.

[0215] Output: The server receives the edited animation data.

[0216] Step 9:

[0217] Server: Saves the edited animation data to the database and retransmits the saved data to the device as needed.

[0218] Input: JSON data of the edited animation sequence received from the device.

[0219] What happens: The server stores the animation data in a database and prepares it for retransmission.

[0220] Output: The saved data is resent to the device.

[0221] As a result, users can efficiently and easily generate and edit high-quality character animations without requiring specialized knowledge. This system also makes it easy to save and reuse data, helping to reduce developer effort and costs.

[0222] (Application example 1)

[0223] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0224] Current animation editing systems make it difficult for users to easily customize the atmosphere and usage scenes of animations, and checking and correcting the generated data is time-consuming. Furthermore, there are limited ways to efficiently provide customized content to users. This makes it difficult to generate and edit animations and content that meet the diverse needs of users.

[0225] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0226] In this invention, the server includes means for allowing a user to select an animation atmosphere and usage scene, means for generating predetermined animation data based on the selected atmosphere and usage scene, means for providing the generated animation data to the user, means for allowing the user to edit the provided animation data, means for saving the edited animation data, and means for automatically generating animations of scenes in videos or virtual reality content based on the atmosphere and usage scene selected by the user. This allows users to easily create, edit, check, and save animations and content that suit their preferences, enabling the provision of customized content that meets a variety of needs.

[0227] "User" means a person who uses the animation system to select an atmosphere and usage scenario, and creates, edits, checks, and saves customized content.

[0228] "Mood" describes the look and feel of your animation or content, and includes characteristics such as casual, formal, or playful.

[0229] "Usage scene" refers to the specific scene or situation in which the animation or content is used, and can be of various types such as everyday life, combat, or fantasy.

[0230] A "means" is a method, process, or device for realizing a specific function within an animation generation and editing system.

[0231] "Animation data" is data that digitally expresses the movements of characters and scenes generated based on the atmosphere and scene in which they are used.

[0232] "Generation" is the process of creating animation data under specific conditions based on user selections.

[0233] "Providing" refers to the act of displaying and transmitting the generated animation data to the user.

[0234] "Editing" is the process in which the user makes changes to the provided animation data and modifies it to a desired form.

[0235] "Storage" refers to the act of recording edited or generated animation data in a database or other storage means and keeping it in a form that can be reused later.

[0236] A "video" is a collection of consecutive image frames, and is a medium for expressing movement and change.

[0237] "Virtual reality content" refers to computer-generated simulated environments or scenes, and is an interactive medium that allows users to feel immersed in the experience.

[0238] This invention is a system that allows users to select the atmosphere and scene of the animation, and then automatically generates, edits, and saves animations of scenes in videos and virtual reality content based on the selected atmosphere and scene. This system can provide customized content that meets the diverse needs of users.

[0239] System configuration

[0240] The system mainly includes the following components:

[0241] User device: Using a smartphone or head-mounted display, it provides an interface for users to select, confirm, and edit.

[0242] Server: A cloud server is used to generate, store, and retransmit data.

[0243] Database: Serves as a repository for storing animation and editing data.

[0244] Program processing overview

[0245] User Input

[0246] On the user's device, the user launches the application and selects the animation mood (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy). The selected information is packaged in JSON format and sent to the server.

[0247] Generate animation

[0248] The server analyzes the input data received from the device and generates appropriate animation data based on the selected mood and usage scenario. Specifically, the server searches for preset motions in the database, selects motions that match the conditions, and combines them to generate a series of animation sequences. The generated animation sequences are sent to the device in JSON format.

[0249] Checking and editing animations

[0250] The generated animation data is displayed on the user's device. The user can use the provided editing tools to fine-tune the motion connections and timeline. The edited animation data is then sent back to the server in JSON format.

[0251] Data storage

[0252] The server receives the edited animation data and stores it in a database. If necessary, the saved data can be sent back to the terminal for final confirmation by the user.

[0253] Hardware and software used

[0254] Hardware: Smartphone, head-mounted display, cloud server

[0255] Software: Python, Requests library, JSON format data

[0256] Specific examples

[0257] When a user opens the app on their smartphone and selects the "casual" mood and the "action" scene, the system automatically generates and displays a casual action scene. The user can then fine-tune the scene and save it once they are satisfied.

[0258] Example prompt for a generative AI model:

[0259] Input prompt for generative AI model: Generate animations for action scenes with a casual atmosphere.

[0260] This allows users to easily create, edit, check, and save animations and content that suit their preferences.

[0261] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0262] Step 1:

[0263] When a user launches the application, the device displays the initial screen, where the user selects the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday, combat, fantasy). This selection information is packaged in JSON format and sent to the server.

[0264] Input: User-selected mood and usage scenario

[0265] Output: Selection information packaged in JSON format

[0266] Step 2:

[0267] The server analyzes the input data received from the device and generates appropriate animation data based on the selected mood and usage scenario. The server searches for preset motions in the database, selects motions that match the conditions, and generates a series of animation sequences. The generated animation sequences are sent to the device in JSON format.

[0268] Input: JSON data of the selection information sent from the terminal

[0269] Output: Animation sequence packaged in JSON format

[0270] Step 3:

[0271] The device receives the animation sequence sent from the server and displays it to the user. The user can then use the provided editing tools to check and fine-tune the animation. Specific editing tasks include smoothing the joins of motions and correcting the timeline. The edited animation data is then sent back to the server in JSON format.

[0272] Input: JSON data of the animation sequence sent from the server

[0273] Output: JSON data of the animation edited by the user

[0274] Step 4:

[0275] The server receives the edited animation data sent from the device and stores it in a database. If necessary, the saved data is sent back to the device so that the user can make a final confirmation. The saved data also includes the editing history, selected atmosphere, and usage scene information.

[0276] Input: JSON data of the edit animation resent from the device

[0277] Output: Animation data stored in a database

[0278] Through these steps, users can easily create, edit, and save animations and content to suit their preferences.

[0279] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0280] This invention combines a system for easily and efficiently creating character animations for VR games with an emotion engine that recognizes the user's emotions. This invention is implemented as follows.

[0281] Receiving User Input

[0282] Terminal

[0283] When a user launches the application, an initial screen appears, providing an animated atmosphere, a usage scenario, and an interface for selecting or recognizing the user's emotions.

[0284] The user selects the preferred atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy).

[0285] Emotion recognition

[0286] Terminal

[0287] As the user makes a selection, the emotion engine recognizes emotions in real time from the user's facial expressions and voice.

[0288] The recognized emotion information is packaged in JSON format along with the mood and usage scene selection information and sent to the server.

[0289] Generate animation

[0290] server

[0291] The data received from the device is analyzed, and appropriate animation data is selected based on the selected atmosphere, usage scene, and recognized emotion.

[0292] The server searches the preset motions in the database and selects a motion that meets the conditions.

[0293] The selected motions are combined to generate a series of animation sequences, including adjusting animation details based on emotional information.

[0294] The generated animation sequence is sent to the device in JSON format.

[0295] Sending Animations

[0296] server

[0297] The generated animation data is sent to the device and reflected to the user in real time as an HTTP response.

[0298] User viewing and editing of animations

[0299] Terminal

[0300] The user can check the animation movement on the screen based on the received animation data.

[0301] The editing tools provided allow users to fine-tune motion joins and timelines, and automatic editing is also possible based on emotions recognized by the emotion engine.

[0302] The edited animation data is sent back to the server in JSON format.

[0303] Saving and resubmitting animations

[0304] server

[0305] Upon receiving the edited animation data, the server stores the data in a database.

[0306] If necessary, the saved data is sent again to the terminal so that the user can make a final confirmation.

[0307] Specific examples

[0308] 1. Receiving user input and recognizing emotions

[0309] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[0310] 2. Device: The device analyzes the user's face using a camera and recognizes their current emotion as "happiness."

[0311] 2. Creating and sending animations

[0312] 3. Server: Based on the received data, search for casual, everyday walking motions and adjust them to match the emotion of "joy."

[0313] 4. Server: Sends the generated animation to the device.

[0314] 3. Check and edit the animation

[0315] 5. Terminal: Display the animation sent.

[0316] 6. User: Fine-tune any parts of the motion where the joints are not smooth.

[0317] 7. Terminal: Re-send the edited animation data.

[0318] 4. Save and resubmit animations

[0319] 8. Server: Receives the edited data and stores it in the database.

[0320] 9. Server: Re-sends the saved data to the device for final confirmation.

[0321] This allows users to easily and efficiently generate and edit character animations for VR games. By utilizing an emotion engine, this system allows for fine-tuning to match the user's emotions, resulting in more natural animations.

[0322] The processing flow will be explained below.

[0323] Step 1:

[0324] User: Launches the application and the interface for selecting mood, usage scenario, and emotion appears on the screen.

[0325] Step 2:

[0326] User: Select "Casual" as the animation atmosphere and "Everyday" as the usage scene.

[0327] Step 3:

[0328] Device: Converts the user-selected mood and usage scenario into JSON format and prepares to launch the emotion engine.

[0329] Step 4:

[0330] Device: The camera captures the user's face and sends the data to the emotion engine, which analyzes the user's facial expressions and recognizes their emotions.

[0331] Step 5:

[0332] Emotion engine: Based on the results of facial expression analysis, the engine recognizes that the user's emotion is "joy." It converts the recognized emotion information into JSON format and returns it to the device.

[0333] Step 6:

[0334] Terminal: The emotion information received from the emotion engine is packaged together with the selected atmosphere and usage scene information and sent to the server.

[0335] Step 7:

[0336] Server: Receives requests sent from the device and extracts the selected mood, usage scenario, and recognized emotion from the JSON data.

[0337] Step 8:

[0338] Server: Accesses the database and searches for appropriate preset motions based on the extracted atmosphere, usage scene, and emotion. Here, motions that fit "Casual," "Everyday," and "Joy" are selected.

[0339] Step 9:

[0340] Server: Combines appropriate motions from the search results to generate a smooth animation sequence, adjusting the details and timing of the animation based on emotional information.

[0341] Step 10:

[0342] Server: Converts the generated animation sequence into JSON format and creates an HTTP response to send to the device.

[0343] Step 11:

[0344] Server: Sends the generated animation data to the device as an HTTP response.

[0345] Step 12:

[0346] Terminal: Receives the response sent from the server, parses the JSON data, and retrieves the animation.

[0347] Step 13:

[0348] Terminal: Displays animation on the screen based on the analyzed animation data.

[0349] Step 14:

[0350] User: Check the displayed animation and fine-tune the motion joints and timeline as necessary.

[0351] Step 15:

[0352] Terminal: Generates updated animation data reflecting the edits made by the user.

[0353] Step 16:

[0354] On the device: Convert the updated animation data into JSON format and create an HTTP request to send it back to the server.

[0355] Step 17:

[0356] On the device: Send the edited animation data to the server using an HTTP POST request.

[0357] Step 18:

[0358] Server: Receives the updated data sent from the device and parses the JSON data again.

[0359] Step 19:

[0360] Server: Saves the parsed update data to the database.

[0361] Step 20:

[0362] Server: Converts the saved update data into JSON format and creates an HTTP response to resend to the device.

[0363] Step 21:

[0364] Server: Re-sends the saved data to the device for final confirmation.

[0365] Step 22:

[0366] Terminal: Receives the final confirmation data sent from the server and displays it on the user interface.

[0367] Step 23:

[0368] User: Perform a final check of the animation and make further adjustments if necessary.

[0369] The above is the processing flow of the animation generation system including emotion recognition.

[0370] Example 2

[0371] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0372] Currently, generating and editing character animations for VR games requires specialized knowledge and advanced technology, which is time-consuming and labor-intensive. Furthermore, it is difficult to naturally change the quality of the animation in response to the user's emotions. Therefore, there is a need for a system that can easily and efficiently generate high-quality animations that reflect the user's emotions.

[0373] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for the user to select the atmosphere and usage scene of the animation, an emotion recognition means for recognizing the user's emotion in real time, and a means for generating animation data based on the selected atmosphere, usage scene, and recognized emotion information. This makes it possible to easily and efficiently generate natural, high-quality animation that reflects the user's emotion.

[0374] "User" refers to a person who uses the system to create and edit animations.

[0375] "Mood" refers to the overall style or feel of the animation, including casual, formal, playful, etc.

[0376] The "usage scene" indicates a specific situation or scene in which the animation is used, and includes, for example, everyday life, battle, fantasy, etc.

[0377] "Emotion recognition means" refers to technology or devices for identifying emotions from a user's facial expressions and voice in real time.

[0378] "Animation data" refers to digital information that expresses the movement of a character, including preset motions and adjustments based on emotional information.

[0379] "Preset motion" refers to standard movements prepared in advance and stored in a database.

[0380] "Editing tools" refers to the interface or tools a user uses to modify, adjust, or change parts of an animation.

[0381] "Storage means" refers to the technology or method for recording edited animation data in a database or other storage.

[0382] "Server" refers to a computer system that processes, stores, and manages data, and also communicates with user terminals.

[0383] "System" refers to a collection of multiple hardware and software components that consistently handle user input, emotion recognition, animation generation, editing, and storage.

[0384] This invention combines a system for easily and efficiently creating character animations for VR games with an emotion engine that recognizes the user's emotions. This invention is implemented as follows.

[0385] Hardware and Software Configuration

[0386] Terminal

[0387] A terminal is a computing device with a standard user interface, including a camera, microphone, display, touch screen, or mouse.

[0388] The emotion recognition engine includes facial expression recognition software (e.g., FaceAPI) and voice recognition software.

[0389] A dedicated editing tool is provided as an interface for editing animations.

[0390] server

[0391] A server is a computer system that communicates with a database and runs programs to process requests from users.

[0392] The database stores preset motions, various animation data, and edited data.

[0393] Process Overview

[0394] Terminal

[0395] 1. The user launches the application and selects the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy) on the initial screen.

[0396] 2. The emotion recognition engine recognizes emotions in real time from the user's facial expressions and voice, and sends them to the server in JSON format along with the selection information.

[0397] server

[0398] 1. The server analyzes the received data and generates appropriate animation data based on the selected atmosphere, usage scene, and recognized emotion.

[0399] 2. Search the database for preset motions, combine motions that meet your criteria to create a series of animation sequences, and fine-tune the details.

[0400] 3. Send the generated animation sequence to the device in JSON format.

[0401] Terminal

[0402] 1. The user reviews the animation received and uses the provided editing tools to fine-tune the motion joins and timeline.

[0403] 2. The edited animation data is sent back to the server in JSON format.

[0404] server

[0405] 1. Save the edited data in the database and resend it to the device if necessary.

[0406] Specific example explanation

[0407] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[0408] 2. Device: The device analyzes the user's face using a camera and recognizes their current emotion as "happiness."

[0409] 3. Server: Based on the received data, it searches for casual, everyday walking motions and adjusts them to match the emotion of "joy."

[0410] 4. Server: Sends the generated animation to the device.

[0411] 5. Device: The submitted animation is displayed and the user can fine-tune any uneven joints in the motion.

[0412] 6. Terminal: The edited animation data is resent, received by the server, and stored in the database.

[0413] Example prompts for generative AI models

[0414] Below are some example prompts to be input to the generative AI model:

[0415] Synopsis

[0416] We have combined a system that allows you to easily create character animations for VR games with a function that recognizes user emotions. We will explain in detail the operating procedure of this system in natural language.

[0417] procedure:

[0418] 1. The user launches the application and selects the mood and usage scenario.

[0419] 2. The device's emotion engine recognizes the user's emotions from their facial expressions and voice in real time and sends the data to the server.

[0420] 3. Based on the data received by the server, the appropriate animation is generated and sent to the device.

[0421] 4. The user checks and edits the animation, and the edited results are sent back to the server.

[0422] 5. The server stores the edited data and resubmits it for final confirmation.

[0423] question:

[0424] Please explain in detail the process of the above system. Please also specify the names of the hardware and software used. Please also include specific examples.

[0425] In this way, this system makes it possible to easily and efficiently generate and edit high-quality animations that reflect the user's emotions.

[0426] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0427] Step 1:

[0428] The user launches the application and selects the animation mood and usage scenario on the initial screen. At this time, the user selects the desired mood (e.g., casual) and usage scenario (e.g., everyday) on the interface using a touchscreen or mouse. The user's selection information is input, and the selection information is saved in the device as output.

[0429] Specific behavior:

[0430] The terminal interface presents the user with options, and the user selects "casual" and "everyday."

[0431] Step 2:

[0432] The device captures the user's facial expressions and voice in real time, which are then analyzed by an emotion recognition engine. A camera and microphone are used for emotion recognition. The input is the captured facial expression and voice data, and the output is the analysis result, which is emotional information such as "happiness."

[0433] Specific behavior:

[0434] The device's camera and microphone capture the user's facial expressions and voice, and an emotion engine (e.g., FaceAPI) recognizes "happiness."

[0435] Step 3:

[0436] The device packages the selection information and emotion information in JSON format and sends it to the server. The input is the user's selection information and emotion information, and a single JSON data is generated as the output and sent to the server.

[0437] Specific behavior:

[0438] The device compiles the selection information and the "joy" emotion information into JSON format and sends it to the server using an HTTP request.

[0439] Step 4:

[0440] The server analyzes the received JSON data and generates animation data based on the selected atmosphere, usage scenario, and emotional information. The received JSON data is the input, and the generated animation data is the output. Preset motions are searched for in the database, motions that match the conditions are selected, and adjustments are made based on the emotional information.

[0441] Specific behavior:

[0442] The server queries the database to find "casual" and "everyday" walking motions and adjusts the "pleasure" of them.

[0443] Step 5:

[0444] The server sends the generated animation data in JSON format to the terminal. The generated animation data is the input, and the JSON format data is sent to the terminal as the output.

[0445] Specific behavior:

[0446] The server converts the generated animation data into JSON format and sends it to the terminal via an HTTP response.

[0447] Step 6:

[0448] The device displays the received animation, and the user can fine-tune the motion joints and timeline using the provided editing tools. The input is the received animation data, and the edited animation data is generated as the output.

[0449] Specific behavior:

[0450] The device plays the animation and the user makes fine adjustments to the timeline.

[0451] Step 7:

[0452] The device repackages the edited animation data into JSON format and sends it to the server.,The input is the edited animation data, and the output is,sent to the server in JSON format.

[0453] Specific behavior:

[0454] The device compiles the edited animation data into JSON format and sends it to the server using an HTTP request.

[0455] Step 8:

[0456] The server receives the edited data and stores it in a database. The input is the received edited animation data, and the output is stored in the database.

[0457] Specific behavior:

[0458] The edited data received by the server is saved in a database and kept as confirmation data.

[0459] This series of steps enables users to easily and efficiently generate and edit character animations for VR games, and also enables natural animations that respond to the user's emotions.

[0460] (Application example 2)

[0461] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0462] Conventional animation generation systems were capable of generating animations based on user input, but they had the problem of being unable to dynamically change the animation to reflect the user's real-time emotions. As a result, the user experience was limited, and it was difficult to achieve natural interactions. In particular, there was a lack of systems in physical stores that could provide guidance and information that took into account the customer's emotions.

[0463] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's emotion using an emotion recognition engine and dynamically adjusting animation based on the recognized emotion, means for the user to select the atmosphere and usage scene of the animation, and means for saving the edited animation data. This makes it possible to provide animation that responds to the user's emotion and to provide personalized guidance in physical stores.

[0464] "Means for the user to select the atmosphere and usage scene of the animation" refers to an interface function that allows the user to select the atmosphere and usage scene of the animation that they prefer.

[0465] The "means for generating predetermined animation data" refers to a function that can appropriately generate pre-stored animation data based on the selected atmosphere and usage scene.

[0466] "Means for providing the generated animation data to the user" refers to a function for displaying or transmitting the generated animation data so that the user can check it.

[0467] "Means for users to edit the provided animation data" refers to editing tools that users can use to modify and adjust the provided animation data.

[0468] "Means for saving edited animation data" refers to a function for saving animation data edited by a user in a database or the like.

[0469] "Means for recognizing a user's emotions using an emotion recognition engine and dynamically adjusting animations based on the recognized emotions" refers to a function that recognizes emotions from a user's facial expressions and voice in real time, and dynamically adjusts the movement and content of animations based on that information.

[0470] "Means for combining multiple preset motions to generate smooth animation" refers to a function that combines multiple preset animation motions to generate continuous, smooth animation.

[0471] "Means for regenerating animation data edited by a user and providing it as saved data" refers to a function for regenerating animation data edited by a user and finally providing it as saved data.

[0472] This invention uses an animation generation system in conjunction with an emotion recognition engine to dynamically adjust animations based on the user's real-time emotions, with the aim of improving the customer experience in brick-and-mortar stores.

[0473] System Configuration

[0474] This system is configured using the following hardware and software.

[0475] Device: Smart glasses or head-mounted display

[0476] Camera: A camera for capturing user facial expressions in real time

[0477] Emotion recognition engine: A software engine that recognizes the user's emotions from facial expressions and voice.

[0478] Server: A server for data processing and animation generation

[0479] Network: A communications infrastructure for data communication between terminals and servers.

[0480] Program processing

[0481] emotion recognition

[0482] The device is used while the user is wearing smart glasses or a head-mounted display. The camera captures the user's facial expressions in real time, and the emotion recognition engine analyzes the user's emotions. The analysis results are sent to the server along with the selected mood and usage scenario.

[0483] Data processing and animation generation

[0484] The server analyzes the received data and generates appropriate animation data based on the selected atmosphere, usage scenario, and recognized emotion. Specifically, it combines multiple preset motions to generate smooth animation. It also dynamically adjusts the details of the animation based on the emotion information. This adjustment enables real-time animation that responds to the user's emotions.

[0485] Providing and editing animations

[0486] The animation data generated from the server is sent to the device and provided to the user on the device. The user can check the provided animation data and use editing tools to correct or adjust the animation as needed. Once editing is complete, the animation data is sent back to the server, where it is finally saved and provided again.

[0487] Specific examples

[0488] For example, when a user is smiling in a new product section during a store introduction or product introduction, the emotion recognition engine will recognize this as "joy." Based on this information, the server will generate a message saying, "Our new product is perfect for your smile!" and send it to the device.

[0489] Example prompts for generative AI models

[0490] "Generate appropriate dynamic guidance messages based on the customer's facial expression recognition data and section information. Provide the following information: emotion recognition results, the customer's current location, and summary information for each section."

[0491] As a result, this system is able to generate animations that respond to the user's emotions and provide personalized store guidance based on those animations.

[0492] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0493] Step 1:

[0494] The device activates smart glasses or a head-mounted display and captures the user's face with a camera. The input is the user's real-time facial expression data, which is captured as a camera image. This allows the device worn by the user to collect basic data for emotion recognition.

[0495] Step 2:

[0496] The emotion recognition engine installed in the device analyzes the captured facial expression data in real time and recognizes the user's emotions. The input is the camera image and the output is the user's emotion (e.g., joy, sadness, surprise, fatigue). The emotion recognition engine uses an algorithm to analyze facial features and identify emotions.

[0497] Step 3:

[0498] The device packages the recognized emotion information, along with the user-selected mood and usage scenario information, in JSON format and sends it to the server. The input is the user's emotion data and selection information, and the output is JSON data containing these. The HTTP protocol is used for transmission.

[0499] Step 4:

[0500] The server analyzes the received JSON data and generates appropriate animation data based on the selected atmosphere, usage scenario, and recognized emotion. The input is the JSON data received from the device, and the output is the generated animation data. Specifically, preset motions are combined and dynamically adjusted based on emotion information.

[0501] Step 5:

[0502] The server sends the generated animation data to the terminal. The input is the generated animation data, and the output is the animation data sent to the terminal. The HTTP protocol is used again for transmission.

[0503] Step 6:

[0504] The device presents the received animation data to the user, who then checks the animation movement. The input is animation data from the server, and the output is the animation displayed on the user's screen. This allows the user to visually check the generated animation.

[0505] Step 7:

[0506] The user can fine-tune the animation using the provided editing tools. The input is the animation displayed on the device and the user's editing operations, and the output is the animation data edited by the user. The editing tools include a function to adjust the smoothness of the motion joints.

[0507] Step 8:

[0508] The device sends the edited animation data back to the server in JSON format. The input is the animation data edited by the user, and the output is the data sent to the server. The HTTP protocol is used for transmission.

[0509] Step 9:

[0510] The server receives the edited animation data and stores it in the database. The input is the edited animation data received from the device, and the output is the data stored in the database. This way, the edited animation is saved for future use.

[0511] Step 10:

[0512] If necessary, the server sends the saved data back to the terminal so that the user can make a final check. The input is the saved animation data, and the output is the data sent to the terminal for final check. This allows the user to check the animation after the final editing is completed.

[0513] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0514] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0515] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0516] [Second embodiment]

[0517] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0518] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0519] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0520] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0521] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0522] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0523] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0524] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0525] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0526] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0527] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0528] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0529] This invention is a system for easily and efficiently creating character animations for VR games, and is implemented as follows.

[0530] Receiving User Input

[0531] Terminal

[0532] When a user launches the application, an initial screen is displayed, which provides an interface for selecting the animation atmosphere and usage scenario.

[0533] The user selects the preferred atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy).

[0534] The selected information is packaged in JSON format and sent to the server.

[0535] Generate animation

[0536] server

[0537] The input data received from the terminal is analyzed, and appropriate animation data is selected based on the selected atmosphere and usage scene.

[0538] The server searches the preset motions in the database and selects a motion that meets the conditions.

[0539] The selected motions are combined to generate a series of animation sequences.

[0540] The generated animation sequence is sent to the device in JSON format.

[0541] Sending Animations

[0542] server

[0543] The generated animation data is sent to the device as an HTTP response and is reflected to the user in real time.

[0544] The server will continue to maintain the connection after sending to allow for editing and resending.

[0545] User viewing and editing of animations

[0546] Terminal

[0547] The user can check the animation movement on the screen based on the received animation data.

[0548] You can use the editing tools provided to fine-tune motion joins and timelines.

[0549] The edited animation data is sent back to the server in JSON format.

[0550] Saving and resubmitting animations

[0551] server

[0552] Upon receiving the edited animation data, the server stores the data in a database.

[0553] If necessary, the saved data is sent again to the terminal so that the user can make a final confirmation.

[0554] Specific examples

[0555] 1. Receiving User Input

[0556] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[0557] 2. Device: Sends selection to server.

[0558] 2. Creating and sending animations

[0559] 3. Server: Searches and generates casual, everyday walking motions based on the received data.

[0560] 4. Server: Sends the generated animation to the device.

[0561] 3. Check and edit the animation

[0562] 5. Terminal: Display the animation sent.

[0563] 6. User: Fine-tune any parts of the motion where the joints are not smooth.

[0564] 7. Terminal: Re-send the edited animation data.

[0565] 4. Save and resubmit animations

[0566] 8. Server: Receives the edited data and stores it in the database.

[0567] 9. Server: Re-sends the saved data to the device for final confirmation.

[0568] This allows users to easily and efficiently generate and edit character animations for VR games, helping to reduce developer workload and costs.

[0569] The processing flow will be explained below.

[0570] Step 1:

[0571] User: Launches the application and an interface appears on the screen for selecting the atmosphere and usage scenario.

[0572] Step 2:

[0573] User: Select "Casual" as the animation atmosphere and "Everyday" as the usage scene.

[0574] Step 3:

[0575] Terminal: Converts the selected information into JSON format and creates an HTTP request to send to the server.

[0576] Step 4:

[0577] On the device: Send the selected atmosphere and usage scene information to the server using an HTTP POST request.

[0578] Step 5:

[0579] Server: Analyzes the request received from the device and extracts the selected mood and usage scenario from the JSON data.

[0580] Step 6:

[0581] Server: Accesses the database and searches for appropriate preset motions based on the extracted atmosphere and usage scene.

[0582] Step 7:

[0583] Server: Selects appropriate motions from the search results and combines multiple motions to generate smooth animation sequences.

[0584] Step 8:

[0585] Server: Converts the generated animation sequence into JSON format and creates an HTTP response to send to the device.

[0586] Step 9:

[0587] Server: Sends the generated animation data to the device as an HTTP response.

[0588] Step 10:

[0589] Terminal: Receives the response from the server, parses the JSON data, and retrieves the animation.

[0590] Step 11:

[0591] Terminal: Displays animation on the screen based on the analyzed animation data.

[0592] Step 12:

[0593] User: Check the displayed animation and fine-tune the motion joints and timeline as necessary.

[0594] Step 13:

[0595] Terminal: Generates updated animation data reflecting the edits made by the user.

[0596] Step 14:

[0597] On the device: Convert the updated animation data into JSON format and create an HTTP request to send it back to the server.

[0598] Step 15:

[0599] On the device: Send the edited animation data to the server using an HTTP POST request.

[0600] Step 16:

[0601] Server: Receives the updated data sent from the device and parses the JSON data again.

[0602] Step 17:

[0603] Server: Saves the parsed update data to the database.

[0604] Step 18:

[0605] Server: Converts the saved update data into JSON format and creates an HTTP response to resend to the device.

[0606] Step 19:

[0607] Server: Re-sends the saved data to the device for final confirmation.

[0608] Step 20:

[0609] Terminal: Receives the final confirmation data from the server and displays it on the user interface.

[0610] Step 21:

[0611] User: Perform a final check of the animation and make further adjustments if necessary.

[0612] Example 1

[0613] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0614] Conventional character animation creation systems for VR games often require users to perform complex operations and require specialized knowledge, making it difficult to create animations efficiently and easily. Another issue is the inconsistency in the quality of the generated animations.

[0615] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0616] In this invention, the server includes means for analyzing input data using a generative AI model and generating an optimal animation sequence, means for transmitting and receiving JSON format data between the terminal and the server, and means for the user to select the atmosphere and usage scene of the animation. This allows users to efficiently and easily create and edit high-quality animations without requiring specialized knowledge, and also makes it possible to easily save and reuse the data.

[0617] "User" means an individual or organization that uses the system to create and edit character animations for VR games.

[0618] "Animation atmosphere" refers to the style and tone of character animation, and can include casual, formal, playful, etc.

[0619] "Usage scene" refers to an element that indicates the situation or background in which the character animation unfolds, and specifically includes everyday life, battle, fantasy, etc.

[0620] The term "means" refers to a specific mechanism or technology that realizes the functions or methods necessary to carry out the present invention.

[0621] "Generative AI models" refer to artificial intelligence algorithms and programs that automatically generate optimal animation sequences based on user input.

[0622] The "JSON format" is a data format used when sending and receiving data between a terminal and a server, and has a lightweight structure that is easy for humans to read.

[0623] "Terminal" refers to a device that a user uses to access the system, and refers to the hardware on which applications are executed.

[0624] "Server" refers to the central processing unit that analyzes the data sent by the user and generates and manages the animation sequences.

[0625] "Preset motion" refers to predefined standard character movements and actions, and is the basic element for generating animation.

[0626] An "animation sequence" refers to a series of character actions created by combining multiple motions.

[0627] "Database" refers to a storage system for saving generated and edited animation data and reusing it as needed.

[0628] "Editing tools" refers to software features that allow a user to adjust and modify the details of an animation.

[0629] The above definitions clearly define the terms that are included in the claims.

[0630] This invention is a system that allows users to easily and efficiently create and edit character animations for VR games. This system features a series of processes for generating animations based on the atmosphere and usage scene selected by the user.

[0631] Receiving User Input

[0632] Device:

[0633] When a user launches the application, an initial screen is displayed, which provides an interface for selecting the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy). Once the user selects the desired option, the information is packaged in JSON format and sent to the server.

[0634] Sending an animation generation request

[0635] server:

[0636] The system receives and analyzes the JSON data sent from the device. A generative AI model is used to analyze this data. Specifically, it searches a database for appropriate preset motions based on the user's selection, and then selects motions that match the conditions.

[0637] Generate animation

[0638] server:

[0639] Multiple preset motions in the database are combined to generate a series of animation sequences, taking into consideration the smooth connection of the selected motions. The generated animation sequence is then converted back to JSON format and prepared for transmission to the device.

[0640] Sending Animations

[0641] server:

[0642] The generated animation data is sent to the terminal as an HTTP response. This transmission is performed in real time, and the connection is maintained afterwards to allow editing and retransmission.

[0643] User viewing and editing of animations

[0644] Device:

[0645] The received animation data is displayed on the screen, allowing the user to check the movement. Furthermore, the user can fine-tune the motion joints and timeline using the provided editing tools. The edited animation data is then sent back to the server in JSON format.

[0646] Saving and resubmitting animations

[0647] server:

[0648] Once the edited animation data is received, it is saved in the database. If necessary, the saved animation data can be sent back to the terminal so that the user can make a final check.

[0649] This allows users to efficiently and easily generate and edit high-quality character animations without requiring specialized knowledge. Furthermore, the system makes it easy to save and reuse data, thereby contributing to reducing developer workload and costs.

[0650] Specific examples

[0651] Receiving User Input

[0652] 1. User: Launch the app and select the atmosphere "Casual" and the usage scene "Everyday."

[0653] 2. Terminal: Sends the selection to the server in JSON format.

[0654] Generate and send animations

[0655] 3. Server: Searches and generates casual, everyday walking motions based on the received data.

[0656] 4. Server: Sends the generated animation to the device in JSON format.

[0657] Checking and editing animations

[0658] 5. Terminal: Display the animation sent.

[0659] 6. User: Fine-tune any parts of the motion where the joins are not smooth.

[0660] 7. Device: The edited animation data is sent back to the server in JSON format.

[0661] Saving and resubmitting animations

[0662] 8. Server: Receives the edited data and stores it in the database.

[0663] 9. Server: Retransmits the saved data to the device for final confirmation.

[0664] Prompt Sentence Examples

[0665] "Please generate character animations for a VR game with a 'casual' atmosphere and an 'everyday' usage scenario. Please focus on walking motions."

[0666] As a result, this system provides users with an intuitive and easy-to-use interface and advanced animation generation functions.

[0667] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0668] Step 1:

[0669] User: When the application is launched, an initial screen appears, where the user selects the animation mood (e.g., casual, formal, playful) and usage scenario (e.g., everyday life, combat, fantasy).

[0670] Input: User-selected mood and usage scene.

[0671] Specific operation: The user selects "casual" and "everyday."

[0672] Output: The selected information is prepared in the terminal in JSON format.

[0673] Step 2:

[0674] Terminal: Sends the selected information to the server.

[0675] Input: Atmosphere and usage scene data packaged in JSON format.

[0676] Specific operation: The device sends JSON data to the server as an HTTP POST request.

[0677] Output: The server receives the JSON data.

[0678] Step 3:

[0679] Server: Parses the received JSON data and uses a generative AI model to set conditions to generate an appropriate animation sequence from the input data.

[0680] Input: JSON data received from the terminal.

[0681] Specific operation: The server parses the JSON data and searches the database for preset motions that correspond to "casual" and "everyday."

[0682] Output: Get multiple preset motions as search results.

[0683] Step 4:

[0684] Server: Selects preset motions that match the search criteria and combines them to generate a smooth animation sequence.

[0685] Input: Multiple preset motions retrieved from the database.

[0686] Specific operation: The server combines selected preset motions to generate a "casual everyday walking scene."

[0687] Output: The generated animation sequence is structured in JSON format.

[0688] Step 5:

[0689] Server: Sends the generated animation data to the terminal.

[0690] Input: JSON data of the generated animation sequence.

[0691] Specific operation: The server sends the animation JSON data to the device as an HTTP response.

[0692] Output: The device receives the animation sequence.

[0693] Step 6:

[0694] Terminal: The received animation data is displayed on the screen so that the user can check the movement.

[0695] Input: JSON data of the animation sequence received from the server.

[0696] Specific operation: The device plays the animation and provides an interface for the user to confirm.

[0697] Output: The user sees the animation visually.

[0698] Step 7:

[0699] User: Use the provided editing tools to fine-tune the animation, adjusting motion joins and timelines.

[0700] Input: Animation displayed on the device.

[0701] Specific operation: The user adjusts the animation joints by dragging and dropping and checks the smoothness.

[0702] Output: The edited animation is prepared in JSON format on the device.

[0703] Step 8:

[0704] Terminal: Send the edited animation data back to the server.

[0705] Input: JSON data of the user-edited animation sequence.

[0706] Specific operation: The device sends the edited JSON data to the server as an HTTP POST request.

[0707] Output: The server receives the edited animation data.

[0708] Step 9:

[0709] Server: Saves the edited animation data to the database and retransmits the saved data to the device as needed.

[0710] Input: JSON data of the edited animation sequence received from the device.

[0711] What happens: The server stores the animation data in a database and prepares it for retransmission.

[0712] Output: The saved data is resent to the device.

[0713] As a result, users can efficiently and easily generate and edit high-quality character animations without requiring specialized knowledge. This system also makes it easy to save and reuse data, helping to reduce developer effort and costs.

[0714] (Application example 1)

[0715] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0716] Current animation editing systems make it difficult for users to easily customize the atmosphere and usage scenes of animations, and checking and correcting the generated data is time-consuming. Furthermore, there are limited ways to efficiently provide customized content to users. This makes it difficult to generate and edit animations and content that meet the diverse needs of users.

[0717] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0718] In this invention, the server includes means for allowing a user to select an animation atmosphere and usage scene, means for generating predetermined animation data based on the selected atmosphere and usage scene, means for providing the generated animation data to the user, means for allowing the user to edit the provided animation data, means for saving the edited animation data, and means for automatically generating animations of scenes in videos or virtual reality content based on the atmosphere and usage scene selected by the user. This allows users to easily create, edit, check, and save animations and content that suit their preferences, enabling the provision of customized content that meets a variety of needs.

[0719] "User" means a person who uses the animation system to select an atmosphere and usage scenario, and creates, edits, checks, and saves customized content.

[0720] "Mood" describes the look and feel of your animation or content, and includes characteristics such as casual, formal, or playful.

[0721] "Usage scene" refers to the specific scene or situation in which the animation or content is used, and can be of various types such as everyday life, combat, or fantasy.

[0722] A "means" is a method, process, or device for realizing a specific function within an animation generation and editing system.

[0723] "Animation data" is data that digitally expresses the movements of characters and scenes generated based on the atmosphere and scene in which they are used.

[0724] "Generation" is the process of creating animation data under specific conditions based on user selections.

[0725] "Providing" refers to the act of displaying and transmitting the generated animation data to the user.

[0726] "Editing" is the process in which the user makes changes to the provided animation data and modifies it to a desired form.

[0727] "Storage" refers to the act of recording edited or generated animation data in a database or other storage means and keeping it in a form that can be reused later.

[0728] A "video" is a collection of consecutive image frames, and is a medium for expressing movement and change.

[0729] "Virtual reality content" refers to computer-generated simulated environments or scenes, and is an interactive medium that allows users to feel immersed in the experience.

[0730] This invention is a system that allows users to select the atmosphere and scene of the animation, and then automatically generates, edits, and saves animations of scenes in videos and virtual reality content based on the selected atmosphere and scene. This system can provide customized content that meets the diverse needs of users.

[0731] System configuration

[0732] The system mainly includes the following components:

[0733] User device: Using a smartphone or head-mounted display, it provides an interface for users to select, confirm, and edit.

[0734] Server: A cloud server is used to generate, store, and retransmit data.

[0735] Database: Serves as a repository for storing animation and editing data.

[0736] Program processing overview

[0737] User Input

[0738] On the user's device, the user launches the application and selects the animation mood (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy). The selected information is packaged in JSON format and sent to the server.

[0739] Generate animation

[0740] The server analyzes the input data received from the device and generates appropriate animation data based on the selected mood and usage scenario. Specifically, the server searches for preset motions in the database, selects motions that match the conditions, and combines them to generate a series of animation sequences. The generated animation sequences are sent to the device in JSON format.

[0741] Checking and editing animations

[0742] The generated animation data is displayed on the user's device. The user can use the provided editing tools to fine-tune the motion connections and timeline. The edited animation data is then sent back to the server in JSON format.

[0743] Data storage

[0744] The server receives the edited animation data and stores it in a database. If necessary, the saved data can be sent back to the terminal for final confirmation by the user.

[0745] Hardware and software used

[0746] Hardware: Smartphone, head-mounted display, cloud server

[0747] Software: Python, Requests library, JSON format data

[0748] Specific examples

[0749] When a user opens the app on their smartphone and selects the "casual" mood and the "action" scene, the system automatically generates and displays a casual action scene. The user can then fine-tune the scene and save it once they are satisfied.

[0750] Example prompt for a generative AI model:

[0751] Input prompt for generative AI model: Generate animations for action scenes with a casual atmosphere.

[0752] This allows users to easily create, edit, check, and save animations and content that suit their preferences.

[0753] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0754] Step 1:

[0755] When a user launches the application, the device displays the initial screen, where the user selects the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday, combat, fantasy). This selection information is packaged in JSON format and sent to the server.

[0756] Input: User-selected mood and usage scenario

[0757] Output: Selection information packaged in JSON format

[0758] Step 2:

[0759] The server analyzes the input data received from the device and generates appropriate animation data based on the selected mood and usage scenario. The server searches for preset motions in the database, selects motions that match the conditions, and generates a series of animation sequences. The generated animation sequences are sent to the device in JSON format.

[0760] Input: JSON data of the selection information sent from the terminal

[0761] Output: Animation sequence packaged in JSON format

[0762] Step 3:

[0763] The device receives the animation sequence sent from the server and displays it to the user. The user can then use the provided editing tools to check and fine-tune the animation. Specific editing tasks include smoothing the joins of motions and correcting the timeline. The edited animation data is then sent back to the server in JSON format.

[0764] Input: JSON data of the animation sequence sent from the server

[0765] Output: JSON data of the animation edited by the user

[0766] Step 4:

[0767] The server receives the edited animation data sent from the device and stores it in a database. If necessary, the saved data is sent back to the device so that the user can make a final confirmation. The saved data also includes the editing history, selected atmosphere, and usage scene information.

[0768] Input: JSON data of the edit animation resent from the device

[0769] Output: Animation data stored in a database

[0770] Through these steps, users can easily create, edit, and save animations and content to suit their preferences.

[0771] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0772] This invention combines a system for easily and efficiently creating character animations for VR games with an emotion engine that recognizes the user's emotions. This invention is implemented as follows.

[0773] Receiving User Input

[0774] Terminal

[0775] When a user launches the application, an initial screen appears, providing an animated atmosphere, a usage scenario, and an interface for selecting or recognizing the user's emotions.

[0776] The user selects the preferred atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy).

[0777] Emotion recognition

[0778] Terminal

[0779] As the user makes a selection, the emotion engine recognizes emotions in real time from the user's facial expressions and voice.

[0780] The recognized emotion information is packaged in JSON format along with the mood and usage scene selection information and sent to the server.

[0781] Generate animation

[0782] server

[0783] The data received from the device is analyzed, and appropriate animation data is selected based on the selected atmosphere, usage scene, and recognized emotion.

[0784] The server searches the preset motions in the database and selects a motion that meets the conditions.

[0785] The selected motions are combined to generate a series of animation sequences, including adjusting animation details based on emotional information.

[0786] The generated animation sequence is sent to the device in JSON format.

[0787] Sending Animations

[0788] server

[0789] The generated animation data is sent to the device and reflected to the user in real time as an HTTP response.

[0790] User viewing and editing of animations

[0791] Terminal

[0792] The user can check the animation movement on the screen based on the received animation data.

[0793] The editing tools provided allow users to fine-tune motion joins and timelines, and automatic editing is also possible based on emotions recognized by the emotion engine.

[0794] The edited animation data is sent back to the server in JSON format.

[0795] Saving and resubmitting animations

[0796] server

[0797] Upon receiving the edited animation data, the server stores the data in a database.

[0798] If necessary, the saved data is sent again to the terminal so that the user can make a final confirmation.

[0799] Specific examples

[0800] 1. Receiving user input and recognizing emotions

[0801] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[0802] 2. Device: The device analyzes the user's face using a camera and recognizes their current emotion as "happiness."

[0803] 2. Creating and sending animations

[0804] 3. Server: Based on the received data, search for casual, everyday walking motions and adjust them to match the emotion of "joy."

[0805] 4. Server: Sends the generated animation to the device.

[0806] 3. Check and edit the animation

[0807] 5. Terminal: Display the animation sent.

[0808] 6. User: Fine-tune any parts of the motion where the joints are not smooth.

[0809] 7. Terminal: Re-send the edited animation data.

[0810] 4. Save and resubmit animations

[0811] 8. Server: Receives the edited data and stores it in the database.

[0812] 9. Server: Re-sends the saved data to the device for final confirmation.

[0813] This allows users to easily and efficiently generate and edit character animations for VR games. By utilizing an emotion engine, this system allows for fine-tuning to match the user's emotions, resulting in more natural animations.

[0814] The processing flow will be explained below.

[0815] Step 1:

[0816] User: Launches the application and the interface for selecting mood, usage scenario, and emotion appears on the screen.

[0817] Step 2:

[0818] User: Select "Casual" as the animation atmosphere and "Everyday" as the usage scene.

[0819] Step 3:

[0820] Device: Converts the user-selected mood and usage scenario into JSON format and prepares to launch the emotion engine.

[0821] Step 4:

[0822] Device: The camera captures the user's face and sends the data to the emotion engine, which analyzes the user's facial expressions and recognizes their emotions.

[0823] Step 5:

[0824] Emotion engine: Based on the results of facial expression analysis, the engine recognizes that the user's emotion is "joy." It converts the recognized emotion information into JSON format and returns it to the device.

[0825] Step 6:

[0826] Terminal: The emotion information received from the emotion engine is packaged together with the selected atmosphere and usage scene information and sent to the server.

[0827] Step 7:

[0828] Server: Receives requests sent from the device and extracts the selected mood, usage scenario, and recognized emotion from the JSON data.

[0829] Step 8:

[0830] Server: Accesses the database and searches for appropriate preset motions based on the extracted atmosphere, usage scene, and emotion. Here, motions that fit "Casual," "Everyday," and "Joy" are selected.

[0831] Step 9:

[0832] Server: Combines appropriate motions from the search results to generate a smooth animation sequence, adjusting the details and timing of the animation based on emotional information.

[0833] Step 10:

[0834] Server: Converts the generated animation sequence into JSON format and creates an HTTP response to send to the device.

[0835] Step 11:

[0836] Server: Sends the generated animation data to the device as an HTTP response.

[0837] Step 12:

[0838] Terminal: Receives the response sent from the server, parses the JSON data, and retrieves the animation.

[0839] Step 13:

[0840] Terminal: Displays animation on the screen based on the analyzed animation data.

[0841] Step 14:

[0842] User: Check the displayed animation and fine-tune the motion joints and timeline as necessary.

[0843] Step 15:

[0844] Terminal: Generates updated animation data reflecting the edits made by the user.

[0845] Step 16:

[0846] On the device: Convert the updated animation data into JSON format and create an HTTP request to send it back to the server.

[0847] Step 17:

[0848] On the device: Send the edited animation data to the server using an HTTP POST request.

[0849] Step 18:

[0850] Server: Receives the updated data sent from the device and parses the JSON data again.

[0851] Step 19:

[0852] Server: Saves the parsed update data to the database.

[0853] Step 20:

[0854] Server: Converts the saved update data into JSON format and creates an HTTP response to resend to the device.

[0855] Step 21:

[0856] Server: Re-sends the saved data to the device for final confirmation.

[0857] Step 22:

[0858] Terminal: Receives the final confirmation data sent from the server and displays it on the user interface.

[0859] Step 23:

[0860] User: Perform a final check of the animation and make further adjustments if necessary.

[0861] The above is the processing flow of the animation generation system including emotion recognition.

[0862] Example 2

[0863] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0864] Currently, generating and editing character animations for VR games requires specialized knowledge and advanced technology, which is time-consuming and labor-intensive. Furthermore, it is difficult to naturally change the quality of the animation in response to the user's emotions. Therefore, there is a need for a system that can easily and efficiently generate high-quality animations that reflect the user's emotions.

[0865] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for the user to select the atmosphere and usage scene of the animation, an emotion recognition means for recognizing the user's emotion in real time, and a means for generating animation data based on the selected atmosphere, usage scene, and recognized emotion information. This makes it possible to easily and efficiently generate natural, high-quality animation that reflects the user's emotion.

[0866] "User" refers to a person who uses the system to create and edit animations.

[0867] "Mood" refers to the overall style or feel of the animation, including casual, formal, playful, etc.

[0868] The "usage scene" indicates a specific situation or scene in which the animation is used, and includes, for example, everyday life, battle, fantasy, etc.

[0869] "Emotion recognition means" refers to technology or devices for identifying emotions from a user's facial expressions and voice in real time.

[0870] "Animation data" refers to digital information that expresses the movement of a character, including preset motions and adjustments based on emotional information.

[0871] "Preset motion" refers to standard movements prepared in advance and stored in a database.

[0872] "Editing tools" refers to the interface or tools a user uses to modify, adjust, or change parts of an animation.

[0873] "Storage means" refers to the technology or method for recording edited animation data in a database or other storage.

[0874] "Server" refers to a computer system that processes, stores, and manages data, and also communicates with user terminals.

[0875] "System" refers to a collection of multiple hardware and software components that consistently handle user input, emotion recognition, animation generation, editing, and storage.

[0876] This invention combines a system for easily and efficiently creating character animations for VR games with an emotion engine that recognizes the user's emotions. This invention is implemented as follows.

[0877] Hardware and Software Configuration

[0878] Terminal

[0879] A terminal is a computing device with a standard user interface, including a camera, microphone, display, touch screen, or mouse.

[0880] The emotion recognition engine includes facial expression recognition software (e.g., FaceAPI) and voice recognition software.

[0881] A dedicated editing tool is provided as an interface for editing animations.

[0882] server

[0883] A server is a computer system that communicates with a database and runs programs to process requests from users.

[0884] The database stores preset motions, various animation data, and edited data.

[0885] Process Overview

[0886] Terminal

[0887] 1. The user launches the application and selects the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy) on the initial screen.

[0888] 2. The emotion recognition engine recognizes emotions in real time from the user's facial expressions and voice, and sends them to the server in JSON format along with the selection information.

[0889] server

[0890] 1. The server analyzes the received data and generates appropriate animation data based on the selected atmosphere, usage scene, and recognized emotion.

[0891] 2. Search the database for preset motions, combine motions that meet your criteria to create a series of animation sequences, and fine-tune the details.

[0892] 3. Send the generated animation sequence to the device in JSON format.

[0893] Terminal

[0894] 1. The user reviews the animation received and uses the provided editing tools to fine-tune the motion joins and timeline.

[0895] 2. The edited animation data is sent back to the server in JSON format.

[0896] server

[0897] 1. Save the edited data in the database and resend it to the device if necessary.

[0898] Specific example explanation

[0899] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[0900] 2. Device: The device analyzes the user's face using a camera and recognizes their current emotion as "happiness."

[0901] 3. Server: Based on the received data, it searches for casual, everyday walking motions and adjusts them to match the emotion of "joy."

[0902] 4. Server: Sends the generated animation to the device.

[0903] 5. Device: The submitted animation is displayed and the user can fine-tune any uneven joints in the motion.

[0904] 6. Terminal: The edited animation data is resent, received by the server, and stored in the database.

[0905] Example prompts for generative AI models

[0906] Below are some example prompts to be input to the generative AI model:

[0907] Synopsis

[0908] We have combined a system that allows you to easily create character animations for VR games with a function that recognizes user emotions. We will explain in detail the operating procedure of this system in natural language.

[0909] procedure:

[0910] 1. The user launches the application and selects the mood and usage scenario.

[0911] 2. The device's emotion engine recognizes the user's emotions from their facial expressions and voice in real time and sends the data to the server.

[0912] 3. Based on the data received by the server, the appropriate animation is generated and sent to the device.

[0913] 4. The user checks and edits the animation, and the edited results are sent back to the server.

[0914] 5. The server stores the edited data and resubmits it for final confirmation.

[0915] question:

[0916] Please explain in detail the process of the above system. Please also specify the names of the hardware and software used. Please also include specific examples.

[0917] In this way, this system makes it possible to easily and efficiently generate and edit high-quality animations that reflect the user's emotions.

[0918] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0919] Step 1:

[0920] The user launches the application and selects the animation mood and usage scenario on the initial screen. At this time, the user selects the desired mood (e.g., casual) and usage scenario (e.g., everyday) on the interface using a touchscreen or mouse. The user's selection information is input, and the selection information is saved in the device as output.

[0921] Specific behavior:

[0922] The terminal interface presents the user with options, and the user selects "casual" and "everyday."

[0923] Step 2:

[0924] The device captures the user's facial expressions and voice in real time, which are then analyzed by an emotion recognition engine. A camera and microphone are used for emotion recognition. The input is the captured facial expression and voice data, and the output is the analysis result, which is emotional information such as "happiness."

[0925] Specific behavior:

[0926] The device's camera and microphone capture the user's facial expressions and voice, and an emotion engine (e.g., FaceAPI) recognizes "happiness."

[0927] Step 3:

[0928] The device packages the selection information and emotion information in JSON format and sends it to the server. The input is the user's selection information and emotion information, and a single JSON data is generated as the output and sent to the server.

[0929] Specific behavior:

[0930] The device compiles the selection information and the "joy" emotion information into JSON format and sends it to the server using an HTTP request.

[0931] Step 4:

[0932] The server analyzes the received JSON data and generates animation data based on the selected atmosphere, usage scenario, and emotional information. The received JSON data is the input, and the generated animation data is the output. Preset motions are searched for in the database, motions that match the conditions are selected, and adjustments are made based on the emotional information.

[0933] Specific behavior:

[0934] The server queries the database to find "casual" and "everyday" walking motions and adjusts the "pleasure" of them.

[0935] Step 5:

[0936] The server sends the generated animation data in JSON format to the terminal. The generated animation data is the input, and the JSON format data is sent to the terminal as the output.

[0937] Specific behavior:

[0938] The server converts the generated animation data into JSON format and sends it to the terminal via an HTTP response.

[0939] Step 6:

[0940] The device displays the received animation, and the user can fine-tune the motion joints and timeline using the provided editing tools. The input is the received animation data, and the edited animation data is generated as the output.

[0941] Specific behavior:

[0942] The device plays the animation and the user makes fine adjustments to the timeline.

[0943] Step 7:

[0944] The device repackages the edited animation data into JSON format and sends it to the server.,The input is the edited animation data, and the output is,sent to the server in JSON format.

[0945] Specific behavior:

[0946] The device compiles the edited animation data into JSON format and sends it to the server using an HTTP request.

[0947] Step 8:

[0948] The server receives the edited data and stores it in a database. The input is the received edited animation data, and the output is stored in the database.

[0949] Specific behavior:

[0950] The edited data received by the server is saved in a database and kept as confirmation data.

[0951] This series of steps enables users to easily and efficiently generate and edit character animations for VR games, and also enables natural animations that respond to the user's emotions.

[0952] (Application example 2)

[0953] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0954] Conventional animation generation systems were capable of generating animations based on user input, but they had the problem of being unable to dynamically change the animation to reflect the user's real-time emotions. As a result, the user experience was limited, and it was difficult to achieve natural interactions. In particular, there was a lack of systems in physical stores that could provide guidance and information that took into account the customer's emotions.

[0955] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's emotion using an emotion recognition engine and dynamically adjusting animation based on the recognized emotion, means for the user to select the atmosphere and usage scene of the animation, and means for saving the edited animation data. This makes it possible to provide animation that responds to the user's emotion and to provide personalized guidance in physical stores.

[0956] "Means for the user to select the atmosphere and usage scene of the animation" refers to an interface function that allows the user to select the atmosphere and usage scene of the animation that they prefer.

[0957] The "means for generating predetermined animation data" refers to a function that can appropriately generate pre-stored animation data based on the selected atmosphere and usage scene.

[0958] "Means for providing the generated animation data to the user" refers to a function for displaying or transmitting the generated animation data so that the user can check it.

[0959] "Means for users to edit the provided animation data" refers to editing tools that users can use to modify and adjust the provided animation data.

[0960] "Means for saving edited animation data" refers to a function for saving animation data edited by a user in a database or the like.

[0961] "Means for recognizing a user's emotions using an emotion recognition engine and dynamically adjusting animations based on the recognized emotions" refers to a function that recognizes emotions from a user's facial expressions and voice in real time, and dynamically adjusts the movement and content of animations based on that information.

[0962] "Means for combining multiple preset motions to generate smooth animation" refers to a function that combines multiple preset animation motions to generate continuous, smooth animation.

[0963] "Means for regenerating animation data edited by a user and providing it as saved data" refers to a function for regenerating animation data edited by a user and finally providing it as saved data.

[0964] This invention uses an animation generation system in conjunction with an emotion recognition engine to dynamically adjust animations based on the user's real-time emotions, with the aim of improving the customer experience in brick-and-mortar stores.

[0965] System Configuration

[0966] This system is configured using the following hardware and software.

[0967] Device: Smart glasses or head-mounted display

[0968] Camera: A camera for capturing user facial expressions in real time

[0969] Emotion recognition engine: A software engine that recognizes the user's emotions from facial expressions and voice.

[0970] Server: A server for data processing and animation generation

[0971] Network: A communications infrastructure for data communication between terminals and servers.

[0972] Program processing

[0973] emotion recognition

[0974] The device is used while the user is wearing smart glasses or a head-mounted display. The camera captures the user's facial expressions in real time, and the emotion recognition engine analyzes the user's emotions. The analysis results are sent to the server along with the selected mood and usage scenario.

[0975] Data processing and animation generation

[0976] The server analyzes the received data and generates appropriate animation data based on the selected atmosphere, usage scenario, and recognized emotion. Specifically, it combines multiple preset motions to generate smooth animation. It also dynamically adjusts the details of the animation based on the emotion information. This adjustment enables real-time animation that responds to the user's emotions.

[0977] Providing and editing animations

[0978] The animation data generated from the server is sent to the device and provided to the user on the device. The user can check the provided animation data and use editing tools to correct or adjust the animation as needed. Once editing is complete, the animation data is sent back to the server, where it is finally saved and provided again.

[0979] Specific examples

[0980] For example, when a user is smiling in a new product section during a store introduction or product introduction, the emotion recognition engine will recognize this as "joy." Based on this information, the server will generate a message saying, "Our new product is perfect for your smile!" and send it to the device.

[0981] Example prompts for generative AI models

[0982] "Generate appropriate dynamic guidance messages based on the customer's facial expression recognition data and section information. Provide the following information: emotion recognition results, the customer's current location, and summary information for each section."

[0983] As a result, this system is able to generate animations that respond to the user's emotions and provide personalized store guidance based on those animations.

[0984] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0985] Step 1:

[0986] The device activates smart glasses or a head-mounted display and captures the user's face with a camera. The input is the user's real-time facial expression data, which is captured as a camera image. This allows the device worn by the user to collect basic data for emotion recognition.

[0987] Step 2:

[0988] The emotion recognition engine installed in the device analyzes the captured facial expression data in real time and recognizes the user's emotions. The input is the camera image and the output is the user's emotion (e.g., joy, sadness, surprise, fatigue). The emotion recognition engine uses an algorithm to analyze facial features and identify emotions.

[0989] Step 3:

[0990] The device packages the recognized emotion information, along with the user-selected mood and usage scenario information, in JSON format and sends it to the server. The input is the user's emotion data and selection information, and the output is JSON data containing these. The HTTP protocol is used for transmission.

[0991] Step 4:

[0992] The server analyzes the received JSON data and generates appropriate animation data based on the selected atmosphere, usage scenario, and recognized emotion. The input is the JSON data received from the device, and the output is the generated animation data. Specifically, preset motions are combined and dynamically adjusted based on emotion information.

[0993] Step 5:

[0994] The server sends the generated animation data to the terminal. The input is the generated animation data, and the output is the animation data sent to the terminal. The HTTP protocol is used again for transmission.

[0995] Step 6:

[0996] The device presents the received animation data to the user, who then checks the animation movement. The input is animation data from the server, and the output is the animation displayed on the user's screen. This allows the user to visually check the generated animation.

[0997] Step 7:

[0998] The user can fine-tune the animation using the provided editing tools. The input is the animation displayed on the device and the user's editing operations, and the output is the animation data edited by the user. The editing tools include a function to adjust the smoothness of the motion joints.

[0999] Step 8:

[1000] The device sends the edited animation data back to the server in JSON format. The input is the animation data edited by the user, and the output is the data sent to the server. The HTTP protocol is used for transmission.

[1001] Step 9:

[1002] The server receives the edited animation data and stores it in the database. The input is the edited animation data received from the device, and the output is the data stored in the database. This way, the edited animation is saved for future use.

[1003] Step 10:

[1004] If necessary, the server sends the saved data back to the terminal so that the user can make a final check. The input is the saved animation data, and the output is the data sent to the terminal for final check. This allows the user to check the animation after the final editing is completed.

[1005] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1006] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1007] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1008] [Third embodiment]

[1009] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1010] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1011] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1012] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1013] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1014] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1015] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1016] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1017] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1018] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1019] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1020] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1021] This invention is a system for easily and efficiently creating character animations for VR games, and is implemented as follows.

[1022] Receiving User Input

[1023] Terminal

[1024] When a user launches the application, an initial screen is displayed, which provides an interface for selecting the animation atmosphere and usage scenario.

[1025] The user selects the preferred atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy).

[1026] The selected information is packaged in JSON format and sent to the server.

[1027] Generate animation

[1028] server

[1029] The input data received from the terminal is analyzed, and appropriate animation data is selected based on the selected atmosphere and usage scene.

[1030] The server searches the preset motions in the database and selects a motion that meets the conditions.

[1031] The selected motions are combined to generate a series of animation sequences.

[1032] The generated animation sequence is sent to the device in JSON format.

[1033] Sending Animations

[1034] server

[1035] The generated animation data is sent to the device as an HTTP response and is reflected to the user in real time.

[1036] The server will continue to maintain the connection after sending to allow for editing and resending.

[1037] User viewing and editing of animations

[1038] Terminal

[1039] The user can check the animation movement on the screen based on the received animation data.

[1040] You can use the editing tools provided to fine-tune motion joins and timelines.

[1041] The edited animation data is sent back to the server in JSON format.

[1042] Saving and resubmitting animations

[1043] server

[1044] Upon receiving the edited animation data, the server stores the data in a database.

[1045] If necessary, the saved data is sent again to the terminal so that the user can make a final confirmation.

[1046] Specific examples

[1047] 1. Receiving User Input

[1048] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[1049] 2. Device: Sends selection to server.

[1050] 2. Creating and sending animations

[1051] 3. Server: Searches and generates casual, everyday walking motions based on the received data.

[1052] 4. Server: Sends the generated animation to the device.

[1053] 3. Check and edit the animation

[1054] 5. Terminal: Display the animation sent.

[1055] 6. User: Fine-tune any parts of the motion where the joints are not smooth.

[1056] 7. Terminal: Re-send the edited animation data.

[1057] 4. Save and resubmit animations

[1058] 8. Server: Receives the edited data and stores it in the database.

[1059] 9. Server: Re-sends the saved data to the device for final confirmation.

[1060] This allows users to easily and efficiently generate and edit character animations for VR games, helping to reduce developer workload and costs.

[1061] The processing flow will be explained below.

[1062] Step 1:

[1063] User: Launches the application and an interface appears on the screen for selecting the atmosphere and usage scenario.

[1064] Step 2:

[1065] User: Select "Casual" as the animation atmosphere and "Everyday" as the usage scene.

[1066] Step 3:

[1067] Terminal: Converts the selected information into JSON format and creates an HTTP request to send to the server.

[1068] Step 4:

[1069] On the device: Send the selected atmosphere and usage scene information to the server using an HTTP POST request.

[1070] Step 5:

[1071] Server: Analyzes the request received from the device and extracts the selected mood and usage scenario from the JSON data.

[1072] Step 6:

[1073] Server: Accesses the database and searches for appropriate preset motions based on the extracted atmosphere and usage scene.

[1074] Step 7:

[1075] Server: Selects appropriate motions from the search results and combines multiple motions to generate smooth animation sequences.

[1076] Step 8:

[1077] Server: Converts the generated animation sequence into JSON format and creates an HTTP response to send to the device.

[1078] Step 9:

[1079] Server: Sends the generated animation data to the device as an HTTP response.

[1080] Step 10:

[1081] Terminal: Receives the response from the server, parses the JSON data, and retrieves the animation.

[1082] Step 11:

[1083] Terminal: Displays animation on the screen based on the analyzed animation data.

[1084] Step 12:

[1085] User: Check the displayed animation and fine-tune the motion joints and timeline as necessary.

[1086] Step 13:

[1087] Terminal: Generates updated animation data reflecting the edits made by the user.

[1088] Step 14:

[1089] On the device: Convert the updated animation data into JSON format and create an HTTP request to send it back to the server.

[1090] Step 15:

[1091] On the device: Send the edited animation data to the server using an HTTP POST request.

[1092] Step 16:

[1093] Server: Receives the updated data sent from the device and parses the JSON data again.

[1094] Step 17:

[1095] Server: Saves the parsed update data to the database.

[1096] Step 18:

[1097] Server: Converts the saved update data into JSON format and creates an HTTP response to resend to the device.

[1098] Step 19:

[1099] Server: Re-sends the saved data to the device for final confirmation.

[1100] Step 20:

[1101] Terminal: Receives the final confirmation data from the server and displays it on the user interface.

[1102] Step 21:

[1103] User: Perform a final check of the animation and make further adjustments if necessary.

[1104] Example 1

[1105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1106] Conventional character animation creation systems for VR games often require users to perform complex operations and require specialized knowledge, making it difficult to create animations efficiently and easily. Another issue is the inconsistency in the quality of the generated animations.

[1107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1108] In this invention, the server includes means for analyzing input data using a generative AI model and generating an optimal animation sequence, means for transmitting and receiving JSON format data between the terminal and the server, and means for the user to select the atmosphere and usage scene of the animation. This allows users to efficiently and easily create and edit high-quality animations without requiring specialized knowledge, and also makes it possible to easily save and reuse the data.

[1109] "User" means an individual or organization that uses the system to create and edit character animations for VR games.

[1110] "Animation atmosphere" refers to the style and tone of character animation, and can include casual, formal, playful, etc.

[1111] "Usage scene" refers to an element that indicates the situation or background in which the character animation unfolds, and specifically includes everyday life, battle, fantasy, etc.

[1112] The term "means" refers to a specific mechanism or technology that realizes the functions or methods necessary to carry out the present invention.

[1113] "Generative AI models" refer to artificial intelligence algorithms and programs that automatically generate optimal animation sequences based on user input.

[1114] The "JSON format" is a data format used when sending and receiving data between a terminal and a server, and has a lightweight structure that is easy for humans to read.

[1115] "Terminal" refers to a device that a user uses to access the system, and refers to the hardware on which applications are executed.

[1116] "Server" refers to the central processing unit that analyzes the data sent by the user and generates and manages the animation sequences.

[1117] "Preset motion" refers to predefined standard character movements and actions, and is the basic element for generating animation.

[1118] An "animation sequence" refers to a series of character actions created by combining multiple motions.

[1119] "Database" refers to a storage system for saving generated and edited animation data and reusing it as needed.

[1120] "Editing tools" refers to software features that allow a user to adjust and modify the details of an animation.

[1121] The above definitions clearly define the terms that are included in the claims.

[1122] This invention is a system that allows users to easily and efficiently create and edit character animations for VR games. This system features a series of processes for generating animations based on the atmosphere and usage scene selected by the user.

[1123] Receiving User Input

[1124] Device:

[1125] When a user launches the application, an initial screen is displayed, which provides an interface for selecting the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy). Once the user selects the desired option, the information is packaged in JSON format and sent to the server.

[1126] Sending an animation generation request

[1127] server:

[1128] The system receives and analyzes the JSON data sent from the device. A generative AI model is used to analyze this data. Specifically, it searches a database for appropriate preset motions based on the user's selection, and then selects motions that match the conditions.

[1129] Generate animation

[1130] server:

[1131] Multiple preset motions in the database are combined to generate a series of animation sequences, taking into consideration the smooth connection of the selected motions. The generated animation sequence is then converted back to JSON format and prepared for transmission to the device.

[1132] Sending Animations

[1133] server:

[1134] The generated animation data is sent to the terminal as an HTTP response. This transmission is performed in real time, and the connection is maintained afterwards to allow editing and retransmission.

[1135] User viewing and editing of animations

[1136] Device:

[1137] The received animation data is displayed on the screen, allowing the user to check the movement. Furthermore, the user can fine-tune the motion joints and timeline using the provided editing tools. The edited animation data is then sent back to the server in JSON format.

[1138] Saving and resubmitting animations

[1139] server:

[1140] Once the edited animation data is received, it is saved in the database. If necessary, the saved animation data can be sent back to the terminal so that the user can make a final check.

[1141] This allows users to efficiently and easily generate and edit high-quality character animations without requiring specialized knowledge. Furthermore, the system makes it easy to save and reuse data, thereby contributing to reducing developer workload and costs.

[1142] Specific examples

[1143] Receiving User Input

[1144] 1. User: Launch the app and select the atmosphere "Casual" and the usage scene "Everyday."

[1145] 2. Terminal: Sends the selection to the server in JSON format.

[1146] Generate and send animations

[1147] 3. Server: Searches and generates casual, everyday walking motions based on the received data.

[1148] 4. Server: Sends the generated animation to the device in JSON format.

[1149] Checking and editing animations

[1150] 5. Terminal: Display the animation sent.

[1151] 6. User: Fine-tune any parts of the motion where the joins are not smooth.

[1152] 7. Device: The edited animation data is sent back to the server in JSON format.

[1153] Saving and resubmitting animations

[1154] 8. Server: Receives the edited data and stores it in the database.

[1155] 9. Server: Retransmits the saved data to the device for final confirmation.

[1156] Prompt Sentence Examples

[1157] "Please generate character animations for a VR game with a 'casual' atmosphere and an 'everyday' usage scenario. Please focus on walking motions."

[1158] As a result, this system provides users with an intuitive and easy-to-use interface and advanced animation generation functions.

[1159] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1160] Step 1:

[1161] User: When the application is launched, an initial screen appears, where the user selects the animation mood (e.g., casual, formal, playful) and usage scenario (e.g., everyday life, combat, fantasy).

[1162] Input: User-selected mood and usage scene.

[1163] Specific operation: The user selects "casual" and "everyday."

[1164] Output: The selected information is prepared in the terminal in JSON format.

[1165] Step 2:

[1166] Terminal: Sends the selected information to the server.

[1167] Input: Atmosphere and usage scene data packaged in JSON format.

[1168] Specific operation: The device sends JSON data to the server as an HTTP POST request.

[1169] Output: The server receives the JSON data.

[1170] Step 3:

[1171] Server: Parses the received JSON data and uses a generative AI model to set conditions to generate an appropriate animation sequence from the input data.

[1172] Input: JSON data received from the terminal.

[1173] Specific operation: The server parses the JSON data and searches the database for preset motions that correspond to "casual" and "everyday."

[1174] Output: Get multiple preset motions as search results.

[1175] Step 4:

[1176] Server: Selects preset motions that match the search criteria and combines them to generate a smooth animation sequence.

[1177] Input: Multiple preset motions retrieved from the database.

[1178] Specific operation: The server combines selected preset motions to generate a "casual everyday walking scene."

[1179] Output: The generated animation sequence is structured in JSON format.

[1180] Step 5:

[1181] Server: Sends the generated animation data to the terminal.

[1182] Input: JSON data of the generated animation sequence.

[1183] Specific operation: The server sends the animation JSON data to the device as an HTTP response.

[1184] Output: The device receives the animation sequence.

[1185] Step 6:

[1186] Terminal: The received animation data is displayed on the screen so that the user can check the movement.

[1187] Input: JSON data of the animation sequence received from the server.

[1188] Specific operation: The device plays the animation and provides an interface for the user to confirm.

[1189] Output: The user sees the animation visually.

[1190] Step 7:

[1191] User: Use the provided editing tools to fine-tune the animation, adjusting motion joins and timelines.

[1192] Input: Animation displayed on the device.

[1193] Specific operation: The user adjusts the animation joints by dragging and dropping and checks the smoothness.

[1194] Output: The edited animation is prepared in JSON format on the device.

[1195] Step 8:

[1196] Terminal: Send the edited animation data back to the server.

[1197] Input: JSON data of the user-edited animation sequence.

[1198] Specific operation: The device sends the edited JSON data to the server as an HTTP POST request.

[1199] Output: The server receives the edited animation data.

[1200] Step 9:

[1201] Server: Saves the edited animation data to the database and retransmits the saved data to the device as needed.

[1202] Input: JSON data of the edited animation sequence received from the device.

[1203] What happens: The server stores the animation data in a database and prepares it for retransmission.

[1204] Output: The saved data is resent to the device.

[1205] As a result, users can efficiently and easily generate and edit high-quality character animations without requiring specialized knowledge. This system also makes it easy to save and reuse data, helping to reduce developer effort and costs.

[1206] (Application example 1)

[1207] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1208] Current animation editing systems make it difficult for users to easily customize the atmosphere and usage scenes of animations, and checking and correcting the generated data is time-consuming. Furthermore, there are limited ways to efficiently provide customized content to users. This makes it difficult to generate and edit animations and content that meet the diverse needs of users.

[1209] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1210] In this invention, the server includes means for allowing a user to select an animation atmosphere and usage scene, means for generating predetermined animation data based on the selected atmosphere and usage scene, means for providing the generated animation data to the user, means for allowing the user to edit the provided animation data, means for saving the edited animation data, and means for automatically generating animations of scenes in videos or virtual reality content based on the atmosphere and usage scene selected by the user. This allows users to easily create, edit, check, and save animations and content that suit their preferences, enabling the provision of customized content that meets a variety of needs.

[1211] "User" means a person who uses the animation system to select an atmosphere and usage scenario, and creates, edits, checks, and saves customized content.

[1212] "Mood" describes the look and feel of your animation or content, and includes characteristics such as casual, formal, or playful.

[1213] "Usage scene" refers to the specific scene or situation in which the animation or content is used, and can be of various types such as everyday life, combat, or fantasy.

[1214] A "means" is a method, process, or device for realizing a specific function within an animation generation and editing system.

[1215] "Animation data" is data that digitally expresses the movements of characters and scenes generated based on the atmosphere and scene in which they are used.

[1216] "Generation" is the process of creating animation data under specific conditions based on user selections.

[1217] "Providing" refers to the act of displaying and transmitting the generated animation data to the user.

[1218] "Editing" is the process in which the user makes changes to the provided animation data and modifies it to a desired form.

[1219] "Storage" refers to the act of recording edited or generated animation data in a database or other storage means and keeping it in a form that can be reused later.

[1220] A "video" is a collection of consecutive image frames, and is a medium for expressing movement and change.

[1221] "Virtual reality content" refers to computer-generated simulated environments or scenes, and is an interactive medium that allows users to feel immersed in the experience.

[1222] This invention is a system that allows users to select the atmosphere and scene of the animation, and then automatically generates, edits, and saves animations of scenes in videos and virtual reality content based on the selected atmosphere and scene. This system can provide customized content that meets the diverse needs of users.

[1223] System configuration

[1224] The system mainly includes the following components:

[1225] User device: Using a smartphone or head-mounted display, it provides an interface for users to select, confirm, and edit.

[1226] Server: A cloud server is used to generate, store, and retransmit data.

[1227] Database: Serves as a repository for storing animation and editing data.

[1228] Program processing overview

[1229] User Input

[1230] On the user's device, the user launches the application and selects the animation mood (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy). The selected information is packaged in JSON format and sent to the server.

[1231] Generate animation

[1232] The server analyzes the input data received from the device and generates appropriate animation data based on the selected mood and usage scenario. Specifically, the server searches for preset motions in the database, selects motions that match the conditions, and combines them to generate a series of animation sequences. The generated animation sequences are sent to the device in JSON format.

[1233] Checking and editing animations

[1234] The generated animation data is displayed on the user's device. The user can use the provided editing tools to fine-tune the motion connections and timeline. The edited animation data is then sent back to the server in JSON format.

[1235] Data storage

[1236] The server receives the edited animation data and stores it in a database. If necessary, the saved data can be sent back to the terminal for final confirmation by the user.

[1237] Hardware and software used

[1238] Hardware: Smartphone, head-mounted display, cloud server

[1239] Software: Python, Requests library, JSON format data

[1240] Specific examples

[1241] When a user opens the app on their smartphone and selects the "casual" mood and the "action" scene, the system automatically generates and displays a casual action scene. The user can then fine-tune the scene and save it once they are satisfied.

[1242] Example prompt for a generative AI model:

[1243] Input prompt for generative AI model: Generate animations for action scenes with a casual atmosphere.

[1244] This allows users to easily create, edit, check, and save animations and content that suit their preferences.

[1245] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1246] Step 1:

[1247] When a user launches the application, the device displays the initial screen, where the user selects the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday, combat, fantasy). This selection information is packaged in JSON format and sent to the server.

[1248] Input: User-selected mood and usage scenario

[1249] Output: Selection information packaged in JSON format

[1250] Step 2:

[1251] The server analyzes the input data received from the device and generates appropriate animation data based on the selected mood and usage scenario. The server searches for preset motions in the database, selects motions that match the conditions, and generates a series of animation sequences. The generated animation sequences are sent to the device in JSON format.

[1252] Input: JSON data of the selection information sent from the terminal

[1253] Output: Animation sequence packaged in JSON format

[1254] Step 3:

[1255] The device receives the animation sequence sent from the server and displays it to the user. The user can then use the provided editing tools to check and fine-tune the animation. Specific editing tasks include smoothing the joins of motions and correcting the timeline. The edited animation data is then sent back to the server in JSON format.

[1256] Input: JSON data of the animation sequence sent from the server

[1257] Output: JSON data of the animation edited by the user

[1258] Step 4:

[1259] The server receives the edited animation data sent from the device and stores it in a database. If necessary, the saved data is sent back to the device so that the user can make a final confirmation. The saved data also includes the editing history, selected atmosphere, and usage scene information.

[1260] Input: JSON data of the edit animation resent from the device

[1261] Output: Animation data stored in a database

[1262] Through these steps, users can easily create, edit, and save animations and content to suit their preferences.

[1263] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1264] This invention combines a system for easily and efficiently creating character animations for VR games with an emotion engine that recognizes the user's emotions. This invention is implemented as follows.

[1265] Receiving User Input

[1266] Terminal

[1267] When a user launches the application, an initial screen appears, providing an animated atmosphere, a usage scenario, and an interface for selecting or recognizing the user's emotions.

[1268] The user selects the preferred atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy).

[1269] Emotion recognition

[1270] Terminal

[1271] As the user makes a selection, the emotion engine recognizes emotions in real time from the user's facial expressions and voice.

[1272] The recognized emotion information is packaged in JSON format along with the mood and usage scene selection information and sent to the server.

[1273] Generate animation

[1274] server

[1275] The data received from the device is analyzed, and appropriate animation data is selected based on the selected atmosphere, usage scene, and recognized emotion.

[1276] The server searches the preset motions in the database and selects a motion that meets the conditions.

[1277] The selected motions are combined to generate a series of animation sequences, including adjusting animation details based on emotional information.

[1278] The generated animation sequence is sent to the device in JSON format.

[1279] Sending Animations

[1280] server

[1281] The generated animation data is sent to the device and reflected to the user in real time as an HTTP response.

[1282] User viewing and editing of animations

[1283] Terminal

[1284] The user can check the animation movement on the screen based on the received animation data.

[1285] The editing tools provided allow users to fine-tune motion joins and timelines, and automatic editing is also possible based on emotions recognized by the emotion engine.

[1286] The edited animation data is sent back to the server in JSON format.

[1287] Saving and resubmitting animations

[1288] server

[1289] Upon receiving the edited animation data, the server stores the data in a database.

[1290] If necessary, the saved data is sent again to the terminal so that the user can make a final confirmation.

[1291] Specific examples

[1292] 1. Receiving user input and recognizing emotions

[1293] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[1294] 2. Device: The device analyzes the user's face using a camera and recognizes their current emotion as "happiness."

[1295] 2. Creating and sending animations

[1296] 3. Server: Based on the received data, search for casual, everyday walking motions and adjust them to match the emotion of "joy."

[1297] 4. Server: Sends the generated animation to the device.

[1298] 3. Check and edit the animation

[1299] 5. Terminal: Display the animation sent.

[1300] 6. User: Fine-tune any parts of the motion where the joints are not smooth.

[1301] 7. Terminal: Re-send the edited animation data.

[1302] 4. Save and resubmit animations

[1303] 8. Server: Receives the edited data and stores it in the database.

[1304] 9. Server: Re-sends the saved data to the device for final confirmation.

[1305] This allows users to easily and efficiently generate and edit character animations for VR games. By utilizing an emotion engine, this system allows for fine-tuning to match the user's emotions, resulting in more natural animations.

[1306] The processing flow will be explained below.

[1307] Step 1:

[1308] User: Launches the application and the interface for selecting mood, usage scenario, and emotion appears on the screen.

[1309] Step 2:

[1310] User: Select "Casual" as the animation atmosphere and "Everyday" as the usage scene.

[1311] Step 3:

[1312] Device: Converts the user-selected mood and usage scenario into JSON format and prepares to launch the emotion engine.

[1313] Step 4:

[1314] Device: The camera captures the user's face and sends the data to the emotion engine, which analyzes the user's facial expressions and recognizes their emotions.

[1315] Step 5:

[1316] Emotion engine: Based on the results of facial expression analysis, the engine recognizes that the user's emotion is "joy." It converts the recognized emotion information into JSON format and returns it to the device.

[1317] Step 6:

[1318] Terminal: The emotion information received from the emotion engine is packaged together with the selected atmosphere and usage scene information and sent to the server.

[1319] Step 7:

[1320] Server: Receives requests sent from the device and extracts the selected mood, usage scenario, and recognized emotion from the JSON data.

[1321] Step 8:

[1322] Server: Accesses the database and searches for appropriate preset motions based on the extracted atmosphere, usage scene, and emotion. Here, motions that fit "Casual," "Everyday," and "Joy" are selected.

[1323] Step 9:

[1324] Server: Combines appropriate motions from the search results to generate a smooth animation sequence, adjusting the details and timing of the animation based on emotional information.

[1325] Step 10:

[1326] Server: Converts the generated animation sequence into JSON format and creates an HTTP response to send to the device.

[1327] Step 11:

[1328] Server: Sends the generated animation data to the device as an HTTP response.

[1329] Step 12:

[1330] Terminal: Receives the response sent from the server, parses the JSON data, and retrieves the animation.

[1331] Step 13:

[1332] Terminal: Displays animation on the screen based on the analyzed animation data.

[1333] Step 14:

[1334] User: Check the displayed animation and fine-tune the motion joints and timeline as necessary.

[1335] Step 15:

[1336] Terminal: Generates updated animation data reflecting the edits made by the user.

[1337] Step 16:

[1338] On the device: Convert the updated animation data into JSON format and create an HTTP request to send it back to the server.

[1339] Step 17:

[1340] On the device: Send the edited animation data to the server using an HTTP POST request.

[1341] Step 18:

[1342] Server: Receives the updated data sent from the device and parses the JSON data again.

[1343] Step 19:

[1344] Server: Saves the parsed update data to the database.

[1345] Step 20:

[1346] Server: Converts the saved update data into JSON format and creates an HTTP response to resend to the device.

[1347] Step 21:

[1348] Server: Re-sends the saved data to the device for final confirmation.

[1349] Step 22:

[1350] Terminal: Receives the final confirmation data sent from the server and displays it on the user interface.

[1351] Step 23:

[1352] User: Perform a final check of the animation and make further adjustments if necessary.

[1353] The above is the processing flow of the animation generation system including emotion recognition.

[1354] Example 2

[1355] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1356] Currently, generating and editing character animations for VR games requires specialized knowledge and advanced technology, which is time-consuming and labor-intensive. Furthermore, it is difficult to naturally change the quality of the animation in response to the user's emotions. Therefore, there is a need for a system that can easily and efficiently generate high-quality animations that reflect the user's emotions.

[1357] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for the user to select the atmosphere and usage scene of the animation, an emotion recognition means for recognizing the user's emotion in real time, and a means for generating animation data based on the selected atmosphere, usage scene, and recognized emotion information. This makes it possible to easily and efficiently generate natural, high-quality animation that reflects the user's emotion.

[1358] "User" refers to a person who uses the system to create and edit animations.

[1359] "Mood" refers to the overall style or feel of the animation, including casual, formal, playful, etc.

[1360] The "usage scene" indicates a specific situation or scene in which the animation is used, and includes, for example, everyday life, battle, fantasy, etc.

[1361] "Emotion recognition means" refers to technology or devices for identifying emotions from a user's facial expressions and voice in real time.

[1362] "Animation data" refers to digital information that expresses the movement of a character, including preset motions and adjustments based on emotional information.

[1363] "Preset motion" refers to standard movements prepared in advance and stored in a database.

[1364] "Editing tools" refers to the interface or tools a user uses to modify, adjust, or change parts of an animation.

[1365] "Storage means" refers to the technology or method for recording edited animation data in a database or other storage.

[1366] "Server" refers to a computer system that processes, stores, and manages data, and also communicates with user terminals.

[1367] "System" refers to a collection of multiple hardware and software components that consistently handle user input, emotion recognition, animation generation, editing, and storage.

[1368] This invention combines a system for easily and efficiently creating character animations for VR games with an emotion engine that recognizes the user's emotions. This invention is implemented as follows.

[1369] Hardware and Software Configuration

[1370] Terminal

[1371] A terminal is a computing device with a standard user interface, including a camera, microphone, display, touch screen, or mouse.

[1372] The emotion recognition engine includes facial expression recognition software (e.g., FaceAPI) and voice recognition software.

[1373] A dedicated editing tool is provided as an interface for editing animations.

[1374] server

[1375] A server is a computer system that communicates with a database and runs programs to process requests from users.

[1376] The database stores preset motions, various animation data, and edited data.

[1377] Process Overview

[1378] Terminal

[1379] 1. The user launches the application and selects the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy) on the initial screen.

[1380] 2. The emotion recognition engine recognizes emotions in real time from the user's facial expressions and voice, and sends them to the server in JSON format along with the selection information.

[1381] server

[1382] 1. The server analyzes the received data and generates appropriate animation data based on the selected atmosphere, usage scene, and recognized emotion.

[1383] 2. Search the database for preset motions, combine motions that meet your criteria to create a series of animation sequences, and fine-tune the details.

[1384] 3. Send the generated animation sequence to the device in JSON format.

[1385] Terminal

[1386] 1. The user reviews the animation received and uses the provided editing tools to fine-tune the motion joins and timeline.

[1387] 2. The edited animation data is sent back to the server in JSON format.

[1388] server

[1389] 1. Save the edited data in the database and resend it to the device if necessary.

[1390] Specific example explanation

[1391] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[1392] 2. Device: The device analyzes the user's face using a camera and recognizes their current emotion as "happiness."

[1393] 3. Server: Based on the received data, it searches for casual, everyday walking motions and adjusts them to match the emotion of "joy."

[1394] 4. Server: Sends the generated animation to the device.

[1395] 5. Device: The submitted animation is displayed and the user can fine-tune any uneven joints in the motion.

[1396] 6. Terminal: The edited animation data is resent, received by the server, and stored in the database.

[1397] Example prompts for generative AI models

[1398] Below are some example prompts to be input to the generative AI model:

[1399] Synopsis

[1400] We have combined a system that allows you to easily create character animations for VR games with a function that recognizes user emotions. We will explain in detail the operating procedure of this system in natural language.

[1401] procedure:

[1402] 1. The user launches the application and selects the mood and usage scenario.

[1403] 2. The device's emotion engine recognizes the user's emotions from their facial expressions and voice in real time and sends the data to the server.

[1404] 3. Based on the data received by the server, the appropriate animation is generated and sent to the device.

[1405] 4. The user checks and edits the animation, and the edited results are sent back to the server.

[1406] 5. The server stores the edited data and resubmits it for final confirmation.

[1407] question:

[1408] Please explain in detail the process of the above system. Please also specify the names of the hardware and software used. Please also include specific examples.

[1409] In this way, this system makes it possible to easily and efficiently generate and edit high-quality animations that reflect the user's emotions.

[1410] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1411] Step 1:

[1412] The user launches the application and selects the animation mood and usage scenario on the initial screen. At this time, the user selects the desired mood (e.g., casual) and usage scenario (e.g., everyday) on the interface using a touchscreen or mouse. The user's selection information is input, and the selection information is saved in the device as output.

[1413] Specific behavior:

[1414] The terminal interface presents the user with options, and the user selects "casual" and "everyday."

[1415] Step 2:

[1416] The device captures the user's facial expressions and voice in real time, which are then analyzed by an emotion recognition engine. A camera and microphone are used for emotion recognition. The input is the captured facial expression and voice data, and the output is the analysis result, which is emotional information such as "happiness."

[1417] Specific behavior:

[1418] The device's camera and microphone capture the user's facial expressions and voice, and an emotion engine (e.g., FaceAPI) recognizes "happiness."

[1419] Step 3:

[1420] The device packages the selection information and emotion information in JSON format and sends it to the server. The input is the user's selection information and emotion information, and a single JSON data is generated as the output and sent to the server.

[1421] Specific behavior:

[1422] The device compiles the selection information and the "joy" emotion information into JSON format and sends it to the server using an HTTP request.

[1423] Step 4:

[1424] The server analyzes the received JSON data and generates animation data based on the selected atmosphere, usage scenario, and emotional information. The received JSON data is the input, and the generated animation data is the output. Preset motions are searched for in the database, motions that match the conditions are selected, and adjustments are made based on the emotional information.

[1425] Specific behavior:

[1426] The server queries the database to find "casual" and "everyday" walking motions and adjusts the "pleasure" of them.

[1427] Step 5:

[1428] The server sends the generated animation data in JSON format to the terminal. The generated animation data is the input, and the JSON format data is sent to the terminal as the output.

[1429] Specific behavior:

[1430] The server converts the generated animation data into JSON format and sends it to the terminal via an HTTP response.

[1431] Step 6:

[1432] The device displays the received animation, and the user can fine-tune the motion joints and timeline using the provided editing tools. The input is the received animation data, and the edited animation data is generated as the output.

[1433] Specific behavior:

[1434] The device plays the animation and the user makes fine adjustments to the timeline.

[1435] Step 7:

[1436] The device repackages the edited animation data into JSON format and sends it to the server.,The input is the edited animation data, and the output is,sent to the server in JSON format.

[1437] Specific behavior:

[1438] The device compiles the edited animation data into JSON format and sends it to the server using an HTTP request.

[1439] Step 8:

[1440] The server receives the edited data and stores it in a database. The input is the received edited animation data, and the output is stored in the database.

[1441] Specific behavior:

[1442] The edited data received by the server is saved in a database and kept as confirmation data.

[1443] This series of steps enables users to easily and efficiently generate and edit character animations for VR games, and also enables natural animations that respond to the user's emotions.

[1444] (Application example 2)

[1445] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1446] Conventional animation generation systems were capable of generating animations based on user input, but they had the problem of being unable to dynamically change the animation to reflect the user's real-time emotions. As a result, the user experience was limited, and it was difficult to achieve natural interactions. In particular, there was a lack of systems in physical stores that could provide guidance and information that took into account the customer's emotions.

[1447] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's emotion using an emotion recognition engine and dynamically adjusting animation based on the recognized emotion, means for the user to select the atmosphere and usage scene of the animation, and means for saving the edited animation data. This makes it possible to provide animation that responds to the user's emotion and to provide personalized guidance in physical stores.

[1448] "Means for the user to select the atmosphere and usage scene of the animation" refers to an interface function that allows the user to select the atmosphere and usage scene of the animation that they prefer.

[1449] The "means for generating predetermined animation data" refers to a function that can appropriately generate pre-stored animation data based on the selected atmosphere and usage scene.

[1450] "Means for providing the generated animation data to the user" refers to a function for displaying or transmitting the generated animation data so that the user can check it.

[1451] "Means for users to edit the provided animation data" refers to editing tools that users can use to modify and adjust the provided animation data.

[1452] "Means for saving edited animation data" refers to a function for saving animation data edited by a user in a database or the like.

[1453] "Means for recognizing a user's emotions using an emotion recognition engine and dynamically adjusting animations based on the recognized emotions" refers to a function that recognizes emotions from a user's facial expressions and voice in real time, and dynamically adjusts the movement and content of animations based on that information.

[1454] "Means for combining multiple preset motions to generate smooth animation" refers to a function that combines multiple preset animation motions to generate continuous, smooth animation.

[1455] "Means for regenerating animation data edited by a user and providing it as saved data" refers to a function for regenerating animation data edited by a user and finally providing it as saved data.

[1456] This invention uses an animation generation system in conjunction with an emotion recognition engine to dynamically adjust animations based on the user's real-time emotions, with the aim of improving the customer experience in brick-and-mortar stores.

[1457] System Configuration

[1458] This system is configured using the following hardware and software.

[1459] Device: Smart glasses or head-mounted display

[1460] Camera: A camera for capturing user facial expressions in real time

[1461] Emotion recognition engine: A software engine that recognizes the user's emotions from facial expressions and voice.

[1462] Server: A server for data processing and animation generation

[1463] Network: A communications infrastructure for data communication between terminals and servers.

[1464] Program processing

[1465] emotion recognition

[1466] The device is used while the user is wearing smart glasses or a head-mounted display. The camera captures the user's facial expressions in real time, and the emotion recognition engine analyzes the user's emotions. The analysis results are sent to the server along with the selected mood and usage scenario.

[1467] Data processing and animation generation

[1468] The server analyzes the received data and generates appropriate animation data based on the selected atmosphere, usage scenario, and recognized emotion. Specifically, it combines multiple preset motions to generate smooth animation. It also dynamically adjusts the details of the animation based on the emotion information. This adjustment enables real-time animation that responds to the user's emotions.

[1469] Providing and editing animations

[1470] The animation data generated from the server is sent to the device and provided to the user on the device. The user can check the provided animation data and use editing tools to correct or adjust the animation as needed. Once editing is complete, the animation data is sent back to the server, where it is finally saved and provided again.

[1471] Specific examples

[1472] For example, when a user is smiling in a new product section during a store introduction or product introduction, the emotion recognition engine will recognize this as "joy." Based on this information, the server will generate a message saying, "Our new product is perfect for your smile!" and send it to the device.

[1473] Example prompts for generative AI models

[1474] "Generate appropriate dynamic guidance messages based on the customer's facial expression recognition data and section information. Provide the following information: emotion recognition results, the customer's current location, and summary information for each section."

[1475] As a result, this system is able to generate animations that respond to the user's emotions and provide personalized store guidance based on those animations.

[1476] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1477] Step 1:

[1478] The device activates smart glasses or a head-mounted display and captures the user's face with a camera. The input is the user's real-time facial expression data, which is captured as a camera image. This allows the device worn by the user to collect basic data for emotion recognition.

[1479] Step 2:

[1480] The emotion recognition engine installed in the device analyzes the captured facial expression data in real time and recognizes the user's emotions. The input is the camera image and the output is the user's emotion (e.g., joy, sadness, surprise, fatigue). The emotion recognition engine uses an algorithm to analyze facial features and identify emotions.

[1481] Step 3:

[1482] The device packages the recognized emotion information, along with the user-selected mood and usage scenario information, in JSON format and sends it to the server. The input is the user's emotion data and selection information, and the output is JSON data containing these. The HTTP protocol is used for transmission.

[1483] Step 4:

[1484] The server analyzes the received JSON data and generates appropriate animation data based on the selected atmosphere, usage scenario, and recognized emotion. The input is the JSON data received from the device, and the output is the generated animation data. Specifically, preset motions are combined and dynamically adjusted based on emotion information.

[1485] Step 5:

[1486] The server sends the generated animation data to the terminal. The input is the generated animation data, and the output is the animation data sent to the terminal. The HTTP protocol is used again for transmission.

[1487] Step 6:

[1488] The device presents the received animation data to the user, who then checks the animation movement. The input is animation data from the server, and the output is the animation displayed on the user's screen. This allows the user to visually check the generated animation.

[1489] Step 7:

[1490] The user can fine-tune the animation using the provided editing tools. The input is the animation displayed on the device and the user's editing operations, and the output is the animation data edited by the user. The editing tools include a function to adjust the smoothness of the motion joints.

[1491] Step 8:

[1492] The device sends the edited animation data back to the server in JSON format. The input is the animation data edited by the user, and the output is the data sent to the server. The HTTP protocol is used for transmission.

[1493] Step 9:

[1494] The server receives the edited animation data and stores it in the database. The input is the edited animation data received from the device, and the output is the data stored in the database. This way, the edited animation is saved for future use.

[1495] Step 10:

[1496] If necessary, the server sends the saved data back to the terminal so that the user can make a final check. The input is the saved animation data, and the output is the data sent to the terminal for final check. This allows the user to check the animation after the final editing is completed.

[1497] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1498] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1499] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1500] [Fourth embodiment]

[1501] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1502] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1503] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1504] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1505] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1506] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1507] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1508] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1509] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1510] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1511] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1512] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1513] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1514] This invention is a system for easily and efficiently creating character animations for VR games, and is implemented as follows.

[1515] Receiving User Input

[1516] Terminal

[1517] When a user launches the application, an initial screen is displayed, which provides an interface for selecting the animation atmosphere and usage scenario.

[1518] The user selects the preferred atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy).

[1519] The selected information is packaged in JSON format and sent to the server.

[1520] Generate animation

[1521] server

[1522] The input data received from the terminal is analyzed, and appropriate animation data is selected based on the selected atmosphere and usage scene.

[1523] The server searches the preset motions in the database and selects a motion that meets the conditions.

[1524] The selected motions are combined to generate a series of animation sequences.

[1525] The generated animation sequence is sent to the device in JSON format.

[1526] Sending Animations

[1527] server

[1528] The generated animation data is sent to the device as an HTTP response and is reflected to the user in real time.

[1529] The server will continue to maintain the connection after sending to allow for editing and resending.

[1530] User viewing and editing of animations

[1531] Terminal

[1532] The user can check the animation movement on the screen based on the received animation data.

[1533] You can use the editing tools provided to fine-tune motion joins and timelines.

[1534] The edited animation data is sent back to the server in JSON format.

[1535] Saving and resubmitting animations

[1536] server

[1537] Upon receiving the edited animation data, the server stores the data in a database.

[1538] If necessary, the saved data is sent again to the terminal so that the user can make a final confirmation.

[1539] Specific examples

[1540] 1. Receiving User Input

[1541] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[1542] 2. Device: Sends selection to server.

[1543] 2. Creating and sending animations

[1544] 3. Server: Searches and generates casual, everyday walking motions based on the received data.

[1545] 4. Server: Sends the generated animation to the device.

[1546] 3. Check and edit the animation

[1547] 5. Terminal: Display the animation sent.

[1548] 6. User: Fine-tune any parts of the motion where the joints are not smooth.

[1549] 7. Terminal: Re-send the edited animation data.

[1550] 4. Save and resubmit animations

[1551] 8. Server: Receives the edited data and stores it in the database.

[1552] 9. Server: Re-sends the saved data to the device for final confirmation.

[1553] This allows users to easily and efficiently generate and edit character animations for VR games, helping to reduce developer workload and costs.

[1554] The processing flow will be explained below.

[1555] Step 1:

[1556] User: Launches the application and an interface appears on the screen for selecting the atmosphere and usage scenario.

[1557] Step 2:

[1558] User: Select "Casual" as the animation atmosphere and "Everyday" as the usage scene.

[1559] Step 3:

[1560] Terminal: Converts the selected information into JSON format and creates an HTTP request to send to the server.

[1561] Step 4:

[1562] On the device: Send the selected atmosphere and usage scene information to the server using an HTTP POST request.

[1563] Step 5:

[1564] Server: Analyzes the request received from the device and extracts the selected mood and usage scenario from the JSON data.

[1565] Step 6:

[1566] Server: Accesses the database and searches for appropriate preset motions based on the extracted atmosphere and usage scene.

[1567] Step 7:

[1568] Server: Selects appropriate motions from the search results and combines multiple motions to generate smooth animation sequences.

[1569] Step 8:

[1570] Server: Converts the generated animation sequence into JSON format and creates an HTTP response to send to the device.

[1571] Step 9:

[1572] Server: Sends the generated animation data to the device as an HTTP response.

[1573] Step 10:

[1574] Terminal: Receives the response from the server, parses the JSON data, and retrieves the animation.

[1575] Step 11:

[1576] Terminal: Displays animation on the screen based on the analyzed animation data.

[1577] Step 12:

[1578] User: Check the displayed animation and fine-tune the motion joints and timeline as necessary.

[1579] Step 13:

[1580] Terminal: Generates updated animation data reflecting the edits made by the user.

[1581] Step 14:

[1582] On the device: Convert the updated animation data into JSON format and create an HTTP request to send it back to the server.

[1583] Step 15:

[1584] On the device: Send the edited animation data to the server using an HTTP POST request.

[1585] Step 16:

[1586] Server: Receives the updated data sent from the device and parses the JSON data again.

[1587] Step 17:

[1588] Server: Saves the parsed update data to the database.

[1589] Step 18:

[1590] Server: Converts the saved update data into JSON format and creates an HTTP response to resend to the device.

[1591] Step 19:

[1592] Server: Re-sends the saved data to the device for final confirmation.

[1593] Step 20:

[1594] Terminal: Receives the final confirmation data from the server and displays it on the user interface.

[1595] Step 21:

[1596] User: Perform a final check of the animation and make further adjustments if necessary.

[1597] Example 1

[1598] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1599] Conventional character animation creation systems for VR games often require users to perform complex operations and require specialized knowledge, making it difficult to create animations efficiently and easily. Another issue is the inconsistency in the quality of the generated animations.

[1600] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1601] In this invention, the server includes means for analyzing input data using a generative AI model and generating an optimal animation sequence, means for transmitting and receiving JSON format data between the terminal and the server, and means for the user to select the atmosphere and usage scene of the animation. This allows users to efficiently and easily create and edit high-quality animations without requiring specialized knowledge, and also makes it possible to easily save and reuse the data.

[1602] "User" means an individual or organization that uses the system to create and edit character animations for VR games.

[1603] "Animation atmosphere" refers to the style and tone of character animation, and can include casual, formal, playful, etc.

[1604] "Usage scene" refers to an element that indicates the situation or background in which the character animation unfolds, and specifically includes everyday life, battle, fantasy, etc.

[1605] The term "means" refers to a specific mechanism or technology that realizes the functions or methods necessary to carry out the present invention.

[1606] "Generative AI models" refer to artificial intelligence algorithms and programs that automatically generate optimal animation sequences based on user input.

[1607] The "JSON format" is a data format used when sending and receiving data between a terminal and a server, and has a lightweight structure that is easy for humans to read.

[1608] "Terminal" refers to a device that a user uses to access the system, and refers to the hardware on which applications are executed.

[1609] "Server" refers to the central processing unit that analyzes the data sent by the user and generates and manages the animation sequences.

[1610] "Preset motion" refers to predefined standard character movements and actions, and is the basic element for generating animation.

[1611] An "animation sequence" refers to a series of character actions created by combining multiple motions.

[1612] "Database" refers to a storage system for saving generated and edited animation data and reusing it as needed.

[1613] "Editing tools" refers to software features that allow a user to adjust and modify the details of an animation.

[1614] The above definitions clearly define the terms that are included in the claims.

[1615] This invention is a system that allows users to easily and efficiently create and edit character animations for VR games. This system features a series of processes for generating animations based on the atmosphere and usage scene selected by the user.

[1616] Receiving User Input

[1617] Device:

[1618] When a user launches the application, an initial screen is displayed, which provides an interface for selecting the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy). Once the user selects the desired option, the information is packaged in JSON format and sent to the server.

[1619] Sending an animation generation request

[1620] server:

[1621] The system receives and analyzes the JSON data sent from the device. A generative AI model is used to analyze this data. Specifically, it searches a database for appropriate preset motions based on the user's selection, and then selects motions that match the conditions.

[1622] Generate animation

[1623] server:

[1624] Multiple preset motions in the database are combined to generate a series of animation sequences, taking into consideration the smooth connection of the selected motions. The generated animation sequence is then converted back to JSON format and prepared for transmission to the device.

[1625] Sending Animations

[1626] server:

[1627] The generated animation data is sent to the terminal as an HTTP response. This transmission is performed in real time, and the connection is maintained afterwards to allow editing and retransmission.

[1628] User viewing and editing of animations

[1629] Device:

[1630] The received animation data is displayed on the screen, allowing the user to check the movement. Furthermore, the user can fine-tune the motion joints and timeline using the provided editing tools. The edited animation data is then sent back to the server in JSON format.

[1631] Saving and resubmitting animations

[1632] server:

[1633] Once the edited animation data is received, it is saved in the database. If necessary, the saved animation data can be sent back to the terminal so that the user can make a final check.

[1634] This allows users to efficiently and easily generate and edit high-quality character animations without requiring specialized knowledge. Furthermore, the system makes it easy to save and reuse data, thereby contributing to reducing developer workload and costs.

[1635] Specific examples

[1636] Receiving User Input

[1637] 1. User: Launch the app and select the atmosphere "Casual" and the usage scene "Everyday."

[1638] 2. Terminal: Sends the selection to the server in JSON format.

[1639] Generate and send animations

[1640] 3. Server: Searches and generates casual, everyday walking motions based on the received data.

[1641] 4. Server: Sends the generated animation to the device in JSON format.

[1642] Checking and editing animations

[1643] 5. Terminal: Display the animation sent.

[1644] 6. User: Fine-tune any parts of the motion where the joins are not smooth.

[1645] 7. Device: The edited animation data is sent back to the server in JSON format.

[1646] Saving and resubmitting animations

[1647] 8. Server: Receives the edited data and stores it in the database.

[1648] 9. Server: Retransmits the saved data to the device for final confirmation.

[1649] Prompt Sentence Examples

[1650] "Please generate character animations for a VR game with a 'casual' atmosphere and an 'everyday' usage scenario. Please focus on walking motions."

[1651] As a result, this system provides users with an intuitive and easy-to-use interface and advanced animation generation functions.

[1652] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1653] Step 1:

[1654] User: When the application is launched, an initial screen appears, where the user selects the animation mood (e.g., casual, formal, playful) and usage scenario (e.g., everyday life, combat, fantasy).

[1655] Input: User-selected mood and usage scene.

[1656] Specific operation: The user selects "casual" and "everyday."

[1657] Output: The selected information is prepared in the terminal in JSON format.

[1658] Step 2:

[1659] Terminal: Sends the selected information to the server.

[1660] Input: Atmosphere and usage scene data packaged in JSON format.

[1661] Specific operation: The device sends JSON data to the server as an HTTP POST request.

[1662] Output: The server receives the JSON data.

[1663] Step 3:

[1664] Server: Parses the received JSON data and uses a generative AI model to set conditions to generate an appropriate animation sequence from the input data.

[1665] Input: JSON data received from the terminal.

[1666] Specific operation: The server parses the JSON data and searches the database for preset motions that correspond to "casual" and "everyday."

[1667] Output: Get multiple preset motions as search results.

[1668] Step 4:

[1669] Server: Selects preset motions that match the search criteria and combines them to generate a smooth animation sequence.

[1670] Input: Multiple preset motions retrieved from the database.

[1671] Specific operation: The server combines selected preset motions to generate a "casual everyday walking scene."

[1672] Output: The generated animation sequence is structured in JSON format.

[1673] Step 5:

[1674] Server: Sends the generated animation data to the terminal.

[1675] Input: JSON data of the generated animation sequence.

[1676] Specific operation: The server sends the animation JSON data to the device as an HTTP response.

[1677] Output: The device receives the animation sequence.

[1678] Step 6:

[1679] Terminal: The received animation data is displayed on the screen so that the user can check the movement.

[1680] Input: JSON data of the animation sequence received from the server.

[1681] Specific operation: The device plays the animation and provides an interface for the user to confirm.

[1682] Output: The user sees the animation visually.

[1683] Step 7:

[1684] User: Use the provided editing tools to fine-tune the animation, adjusting motion joins and timelines.

[1685] Input: Animation displayed on the device.

[1686] Specific operation: The user adjusts the animation joints by dragging and dropping and checks the smoothness.

[1687] Output: The edited animation is prepared in JSON format on the device.

[1688] Step 8:

[1689] Terminal: Send the edited animation data back to the server.

[1690] Input: JSON data of the user-edited animation sequence.

[1691] Specific operation: The device sends the edited JSON data to the server as an HTTP POST request.

[1692] Output: The server receives the edited animation data.

[1693] Step 9:

[1694] Server: Saves the edited animation data to the database and retransmits the saved data to the device as needed.

[1695] Input: JSON data of the edited animation sequence received from the device.

[1696] What happens: The server stores the animation data in a database and prepares it for retransmission.

[1697] Output: The saved data is resent to the device.

[1698] As a result, users can efficiently and easily generate and edit high-quality character animations without requiring specialized knowledge. This system also makes it easy to save and reuse data, helping to reduce developer effort and costs.

[1699] (Application example 1)

[1700] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1701] Current animation editing systems make it difficult for users to easily customize the atmosphere and usage scenes of animations, and checking and correcting the generated data is time-consuming. Furthermore, there are limited ways to efficiently provide customized content to users. This makes it difficult to generate and edit animations and content that meet the diverse needs of users.

[1702] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1703] In this invention, the server includes means for allowing a user to select an animation atmosphere and usage scene, means for generating predetermined animation data based on the selected atmosphere and usage scene, means for providing the generated animation data to the user, means for allowing the user to edit the provided animation data, means for saving the edited animation data, and means for automatically generating animations of scenes in videos or virtual reality content based on the atmosphere and usage scene selected by the user. This allows users to easily create, edit, check, and save animations and content that suit their preferences, enabling the provision of customized content that meets a variety of needs.

[1704] "User" means a person who uses the animation system to select an atmosphere and usage scenario, and creates, edits, checks, and saves customized content.

[1705] "Mood" describes the look and feel of your animation or content, and includes characteristics such as casual, formal, or playful.

[1706] "Usage scene" refers to the specific scene or situation in which the animation or content is used, and can be of various types such as everyday life, combat, or fantasy.

[1707] A "means" is a method, process, or device for realizing a specific function within an animation generation and editing system.

[1708] "Animation data" is data that digitally expresses the movements of characters and scenes generated based on the atmosphere and scene in which they are used.

[1709] "Generation" is the process of creating animation data under specific conditions based on user selections.

[1710] "Providing" refers to the act of displaying and transmitting the generated animation data to the user.

[1711] "Editing" is the process in which the user makes changes to the provided animation data and modifies it to a desired form.

[1712] "Storage" refers to the act of recording edited or generated animation data in a database or other storage means and keeping it in a form that can be reused later.

[1713] A "video" is a collection of consecutive image frames, and is a medium for expressing movement and change.

[1714] "Virtual reality content" refers to computer-generated simulated environments or scenes, and is an interactive medium that allows users to feel immersed in the experience.

[1715] This invention is a system that allows users to select the atmosphere and scene of the animation, and then automatically generates, edits, and saves animations of scenes in videos and virtual reality content based on the selected atmosphere and scene. This system can provide customized content that meets the diverse needs of users.

[1716] System configuration

[1717] The system mainly includes the following components:

[1718] User device: Using a smartphone or head-mounted display, it provides an interface for users to select, confirm, and edit.

[1719] Server: A cloud server is used to generate, store, and retransmit data.

[1720] Database: Serves as a repository for storing animation and editing data.

[1721] Program processing overview

[1722] User Input

[1723] On the user's device, the user launches the application and selects the animation mood (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy). The selected information is packaged in JSON format and sent to the server.

[1724] Generate animation

[1725] The server analyzes the input data received from the device and generates appropriate animation data based on the selected mood and usage scenario. Specifically, the server searches for preset motions in the database, selects motions that match the conditions, and combines them to generate a series of animation sequences. The generated animation sequences are sent to the device in JSON format.

[1726] Checking and editing animations

[1727] The generated animation data is displayed on the user's device. The user can use the provided editing tools to fine-tune the motion connections and timeline. The edited animation data is then sent back to the server in JSON format.

[1728] Data storage

[1729] The server receives the edited animation data and stores it in a database. If necessary, the saved data can be sent back to the terminal for final confirmation by the user.

[1730] Hardware and software used

[1731] Hardware: Smartphone, head-mounted display, cloud server

[1732] Software: Python, Requests library, JSON format data

[1733] Specific examples

[1734] When a user opens the app on their smartphone and selects the "casual" mood and the "action" scene, the system automatically generates and displays a casual action scene. The user can then fine-tune the scene and save it once they are satisfied.

[1735] Example prompt for a generative AI model:

[1736] Input prompt for generative AI model: Generate animations for action scenes with a casual atmosphere.

[1737] This allows users to easily create, edit, check, and save animations and content that suit their preferences.

[1738] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1739] Step 1:

[1740] When a user launches the application, the device displays the initial screen, where the user selects the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday, combat, fantasy). This selection information is packaged in JSON format and sent to the server.

[1741] Input: User-selected mood and usage scenario

[1742] Output: Selection information packaged in JSON format

[1743] Step 2:

[1744] The server analyzes the input data received from the device and generates appropriate animation data based on the selected mood and usage scenario. The server searches for preset motions in the database, selects motions that match the conditions, and generates a series of animation sequences. The generated animation sequences are sent to the device in JSON format.

[1745] Input: JSON data of the selection information sent from the terminal

[1746] Output: Animation sequence packaged in JSON format

[1747] Step 3:

[1748] The device receives the animation sequence sent from the server and displays it to the user. The user can then use the provided editing tools to check and fine-tune the animation. Specific editing tasks include smoothing the joins of motions and correcting the timeline. The edited animation data is then sent back to the server in JSON format.

[1749] Input: JSON data of the animation sequence sent from the server

[1750] Output: JSON data of the animation edited by the user

[1751] Step 4:

[1752] The server receives the edited animation data sent from the device and stores it in a database. If necessary, the saved data is sent back to the device so that the user can make a final confirmation. The saved data also includes the editing history, selected atmosphere, and usage scene information.

[1753] Input: JSON data of the edit animation resent from the device

[1754] Output: Animation data stored in a database

[1755] Through these steps, users can easily create, edit, and save animations and content to suit their preferences.

[1756] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1757] This invention combines a system for easily and efficiently creating character animations for VR games with an emotion engine that recognizes the user's emotions. This invention is implemented as follows.

[1758] Receiving User Input

[1759] Terminal

[1760] When a user launches the application, an initial screen appears, providing an animated atmosphere, a usage scenario, and an interface for selecting or recognizing the user's emotions.

[1761] The user selects the preferred atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy).

[1762] Emotion recognition

[1763] Terminal

[1764] As the user makes a selection, the emotion engine recognizes emotions in real time from the user's facial expressions and voice.

[1765] The recognized emotion information is packaged in JSON format along with the mood and usage scene selection information and sent to the server.

[1766] Generate animation

[1767] server

[1768] The data received from the device is analyzed, and appropriate animation data is selected based on the selected atmosphere, usage scene, and recognized emotion.

[1769] The server searches the preset motions in the database and selects a motion that meets the conditions.

[1770] The selected motions are combined to generate a series of animation sequences, including adjusting animation details based on emotional information.

[1771] The generated animation sequence is sent to the device in JSON format.

[1772] Sending Animations

[1773] server

[1774] The generated animation data is sent to the device and reflected to the user in real time as an HTTP response.

[1775] User viewing and editing of animations

[1776] Terminal

[1777] The user can check the animation movement on the screen based on the received animation data.

[1778] The editing tools provided allow users to fine-tune motion joins and timelines, and automatic editing is also possible based on emotions recognized by the emotion engine.

[1779] The edited animation data is sent back to the server in JSON format.

[1780] Saving and resubmitting animations

[1781] server

[1782] Upon receiving the edited animation data, the server stores the data in a database.

[1783] If necessary, the saved data is sent again to the terminal so that the user can make a final confirmation.

[1784] Specific examples

[1785] 1. Receiving user input and recognizing emotions

[1786] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[1787] 2. Device: The device analyzes the user's face using a camera and recognizes their current emotion as "happiness."

[1788] 2. Creating and sending animations

[1789] 3. Server: Based on the received data, search for casual, everyday walking motions and adjust them to match the emotion of "joy."

[1790] 4. Server: Sends the generated animation to the device.

[1791] 3. Check and edit the animation

[1792] 5. Terminal: Display the animation sent.

[1793] 6. User: Fine-tune any parts of the motion where the joints are not smooth.

[1794] 7. Terminal: Re-send the edited animation data.

[1795] 4. Save and resubmit animations

[1796] 8. Server: Receives the edited data and stores it in the database.

[1797] 9. Server: Re-sends the saved data to the device for final confirmation.

[1798] This allows users to easily and efficiently generate and edit character animations for VR games. By utilizing an emotion engine, this system allows for fine-tuning to match the user's emotions, resulting in more natural animations.

[1799] The processing flow will be explained below.

[1800] Step 1:

[1801] User: Launches the application and the interface for selecting mood, usage scenario, and emotion appears on the screen.

[1802] Step 2:

[1803] User: Select "Casual" as the animation atmosphere and "Everyday" as the usage scene.

[1804] Step 3:

[1805] Device: Converts the user-selected mood and usage scenario into JSON format and prepares to launch the emotion engine.

[1806] Step 4:

[1807] Device: The camera captures the user's face and sends the data to the emotion engine, which analyzes the user's facial expressions and recognizes their emotions.

[1808] Step 5:

[1809] Emotion engine: Based on the results of facial expression analysis, the engine recognizes that the user's emotion is "joy." It converts the recognized emotion information into JSON format and returns it to the device.

[1810] Step 6:

[1811] Terminal: The emotion information received from the emotion engine is packaged together with the selected atmosphere and usage scene information and sent to the server.

[1812] Step 7:

[1813] Server: Receives requests sent from the device and extracts the selected mood, usage scenario, and recognized emotion from the JSON data.

[1814] Step 8:

[1815] Server: Accesses the database and searches for appropriate preset motions based on the extracted atmosphere, usage scene, and emotion. Here, motions that fit "Casual," "Everyday," and "Joy" are selected.

[1816] Step 9:

[1817] Server: Combines appropriate motions from the search results to generate a smooth animation sequence, adjusting the details and timing of the animation based on emotional information.

[1818] Step 10:

[1819] Server: Converts the generated animation sequence into JSON format and creates an HTTP response to send to the device.

[1820] Step 11:

[1821] Server: Sends the generated animation data to the device as an HTTP response.

[1822] Step 12:

[1823] Terminal: Receives the response sent from the server, parses the JSON data, and retrieves the animation.

[1824] Step 13:

[1825] Terminal: Displays animation on the screen based on the analyzed animation data.

[1826] Step 14:

[1827] User: Check the displayed animation and fine-tune the motion joints and timeline as necessary.

[1828] Step 15:

[1829] Terminal: Generates updated animation data reflecting the edits made by the user.

[1830] Step 16:

[1831] On the device: Convert the updated animation data into JSON format and create an HTTP request to send it back to the server.

[1832] Step 17:

[1833] On the device: Send the edited animation data to the server using an HTTP POST request.

[1834] Step 18:

[1835] Server: Receives the updated data sent from the device and parses the JSON data again.

[1836] Step 19:

[1837] Server: Saves the parsed update data to the database.

[1838] Step 20:

[1839] Server: Converts the saved update data into JSON format and creates an HTTP response to resend to the device.

[1840] Step 21:

[1841] Server: Re-sends the saved data to the device for final confirmation.

[1842] Step 22:

[1843] Terminal: Receives the final confirmation data sent from the server and displays it on the user interface.

[1844] Step 23:

[1845] User: Perform a final check of the animation and make further adjustments if necessary.

[1846] The above is the processing flow of the animation generation system including emotion recognition.

[1847] Example 2

[1848] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1849] Currently, generating and editing character animations for VR games requires specialized knowledge and advanced technology, which is time-consuming and labor-intensive. Furthermore, it is difficult to naturally change the quality of the animation in response to the user's emotions. Therefore, there is a need for a system that can easily and efficiently generate high-quality animations that reflect the user's emotions.

[1850] The specification process by the specification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for the user to select the atmosphere and usage scene of the animation, an emotion recognition means for recognizing the user's emotion in real time, and a means for generating animation data based on the selected atmosphere, usage scene, and recognized emotion information. This makes it possible to easily and efficiently generate natural, high-quality animation that reflects the user's emotion.

[1851] "User" refers to a person who uses the system to create and edit animations.

[1852] "Mood" refers to the overall style or feel of the animation, including casual, formal, playful, etc.

[1853] The "usage scene" indicates a specific situation or scene in which the animation is used, and includes, for example, everyday life, battle, fantasy, etc.

[1854] "Emotion recognition means" refers to technology or devices for identifying emotions from a user's facial expressions and voice in real time.

[1855] "Animation data" refers to digital information that expresses the movement of a character, including preset motions and adjustments based on emotional information.

[1856] "Preset motion" refers to standard movements prepared in advance and stored in a database.

[1857] "Editing tools" refers to the interface or tools a user uses to modify, adjust, or change parts of an animation.

[1858] "Storage means" refers to the technology or method for recording edited animation data in a database or other storage.

[1859] "Server" refers to a computer system that processes, stores, and manages data, and also communicates with user terminals.

[1860] "System" refers to a collection of multiple hardware and software components that consistently handle user input, emotion recognition, animation generation, editing, and storage.

[1861] This invention combines a system for easily and efficiently creating character animations for VR games with an emotion engine that recognizes the user's emotions. This invention is implemented as follows.

[1862] Hardware and Software Configuration

[1863] Terminal

[1864] A terminal is a computing device with a standard user interface, including a camera, microphone, display, touch screen, or mouse.

[1865] The emotion recognition engine includes facial expression recognition software (e.g., FaceAPI) and voice recognition software.

[1866] A dedicated editing tool is provided as an interface for editing animations.

[1867] server

[1868] A server is a computer system that communicates with a database and runs programs to process requests from users.

[1869] The database stores preset motions, various animation data, and edited data.

[1870] Process Overview

[1871] Terminal

[1872] 1. The user launches the application and selects the animation atmosphere (e.g., casual, formal, playful) and usage scene (e.g., everyday life, combat, fantasy) on the initial screen.

[1873] 2. The emotion recognition engine recognizes emotions in real time from the user's facial expressions and voice, and sends them to the server in JSON format along with the selection information.

[1874] server

[1875] 1. The server analyzes the received data and generates appropriate animation data based on the selected atmosphere, usage scene, and recognized emotion.

[1876] 2. Search the database for preset motions, combine motions that meet your criteria to create a series of animation sequences, and fine-tune the details.

[1877] 3. Send the generated animation sequence to the device in JSON format.

[1878] Terminal

[1879] 1. The user reviews the animation received and uses the provided editing tools to fine-tune the motion joins and timeline.

[1880] 2. The edited animation data is sent back to the server in JSON format.

[1881] server

[1882] 1. Save the edited data in the database and resend it to the device if necessary.

[1883] Specific example explanation

[1884] 1. User: Launch the app and select the mood "Casual" and the usage scenario "Everyday."

[1885] 2. Device: The device analyzes the user's face using a camera and recognizes their current emotion as "happiness."

[1886] 3. Server: Based on the received data, it searches for casual, everyday walking motions and adjusts them to match the emotion of "joy."

[1887] 4. Server: Sends the generated animation to the device.

[1888] 5. Device: The submitted animation is displayed and the user can fine-tune any uneven joints in the motion.

[1889] 6. Terminal: The edited animation data is resent, received by the server, and stored in the database.

[1890] Example prompts for generative AI models

[1891] Below are some example prompts to be input to the generative AI model:

[1892] Synopsis

[1893] We have combined a system that allows you to easily create character animations for VR games with a function that recognizes user emotions. We will explain in detail the operating procedure of this system in natural language.

[1894] procedure:

[1895] 1. The user launches the application and selects the mood and usage scenario.

[1896] 2. The device's emotion engine recognizes the user's emotions from their facial expressions and voice in real time and sends the data to the server.

[1897] 3. Based on the data received by the server, the appropriate animation is generated and sent to the device.

[1898] 4. The user checks and edits the animation, and the edited results are sent back to the server.

[1899] 5. The server stores the edited data and resubmits it for final confirmation.

[1900] question:

[1901] Please explain in detail the process of the above system. Please also specify the names of the hardware and software used. Please also include specific examples.

[1902] In this way, this system makes it possible to easily and efficiently generate and edit high-quality animations that reflect the user's emotions.

[1903] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1904] Step 1:

[1905] The user launches the application and selects the animation mood and usage scenario on the initial screen. At this time, the user selects the desired mood (e.g., casual) and usage scenario (e.g., everyday) on the interface using a touchscreen or mouse. The user's selection information is input, and the selection information is saved in the device as output.

[1906] Specific behavior:

[1907] The terminal interface presents the user with options, and the user selects "casual" and "everyday."

[1908] Step 2:

[1909] The device captures the user's facial expressions and voice in real time, which are then analyzed by an emotion recognition engine. A camera and microphone are used for emotion recognition. The input is the captured facial expression and voice data, and the output is the analysis result, which is emotional information such as "happiness."

[1910] Specific behavior:

[1911] The device's camera and microphone capture the user's facial expressions and voice, and an emotion engine (e.g., FaceAPI) recognizes "happiness."

[1912] Step 3:

[1913] The device packages the selection information and emotion information in JSON format and sends it to the server. The input is the user's selection information and emotion information, and a single JSON data is generated as the output and sent to the server.

[1914] Specific behavior:

[1915] The device compiles the selection information and the "joy" emotion information into JSON format and sends it to the server using an HTTP request.

[1916] Step 4:

[1917] The server analyzes the received JSON data and generates animation data based on the selected atmosphere, usage scenario, and emotional information. The received JSON data is the input, and the generated animation data is the output. Preset motions are searched for in the database, motions that match the conditions are selected, and adjustments are made based on the emotional information.

[1918] Specific behavior:

[1919] The server queries the database to find "casual" and "everyday" walking motions and adjusts the "pleasure" of them.

[1920] Step 5:

[1921] The server sends the generated animation data in JSON format to the terminal. The generated animation data is the input, and the JSON format data is sent to the terminal as the output.

[1922] Specific behavior:

[1923] The server converts the generated animation data into JSON format and sends it to the terminal via an HTTP response.

[1924] Step 6:

[1925] The device displays the received animation, and the user can fine-tune the motion joints and timeline using the provided editing tools. The input is the received animation data, and the edited animation data is generated as the output.

[1926] Specific behavior:

[1927] The device plays the animation and the user makes fine adjustments to the timeline.

[1928] Step 7:

[1929] The device repackages the edited animation data into JSON format and sends it to the server.,The input is the edited animation data, and the output is,sent to the server in JSON format.

[1930] Specific behavior:

[1931] The device compiles the edited animation data into JSON format and sends it to the server using an HTTP request.

[1932] Step 8:

[1933] The server receives the edited data and stores it in a database. The input is the received edited animation data, and the output is stored in the database.

[1934] Specific behavior:

[1935] The edited data received by the server is saved in a database and kept as confirmation data.

[1936] This series of steps enables users to easily and efficiently generate and edit character animations for VR games, and also enables natural animations that respond to the user's emotions.

[1937] (Application example 2)

[1938] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1939] Conventional animation generation systems were capable of generating animations based on user input, but they had the problem of being unable to dynamically change the animation to reflect the user's real-time emotions. As a result, the user experience was limited, and it was difficult to achieve natural interactions. In particular, there was a lack of systems in physical stores that could provide guidance and information that took into account the customer's emotions.

[1940] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing a user's emotion using an emotion recognition engine and dynamically adjusting animation based on the recognized emotion, means for the user to select the atmosphere and usage scene of the animation, and means for saving the edited animation data. This makes it possible to provide animation that responds to the user's emotion and to provide personalized guidance in physical stores.

[1941] "Means for the user to select the atmosphere and usage scene of the animation" refers to an interface function that allows the user to select the atmosphere and usage scene of the animation that they prefer.

[1942] The "means for generating predetermined animation data" refers to a function that can appropriately generate pre-stored animation data based on the selected atmosphere and usage scene.

[1943] "Means for providing the generated animation data to the user" refers to a function for displaying or transmitting the generated animation data so that the user can check it.

[1944] "Means for users to edit the provided animation data" refers to editing tools that users can use to modify and adjust the provided animation data.

[1945] "Means for saving edited animation data" refers to a function for saving animation data edited by a user in a database or the like.

[1946] "Means for recognizing a user's emotions using an emotion recognition engine and dynamically adjusting animations based on the recognized emotions" refers to a function that recognizes emotions from a user's facial expressions and voice in real time, and dynamically adjusts the movement and content of animations based on that information.

[1947] "Means for combining multiple preset motions to generate smooth animation" refers to a function that combines multiple preset animation motions to generate continuous, smooth animation.

[1948] "Means for regenerating animation data edited by a user and providing it as saved data" refers to a function for regenerating animation data edited by a user and finally providing it as saved data.

[1949] This invention uses an animation generation system in conjunction with an emotion recognition engine to dynamically adjust animations based on the user's real-time emotions, with the aim of improving the customer experience in brick-and-mortar stores.

[1950] System Configuration

[1951] This system is configured using the following hardware and software.

[1952] Device: Smart glasses or head-mounted display

[1953] Camera: A camera for capturing user facial expressions in real time

[1954] Emotion recognition engine: A software engine that recognizes the user's emotions from facial expressions and voice.

[1955] Server: A server for data processing and animation generation

[1956] Network: A communications infrastructure for data communication between terminals and servers.

[1957] Program processing

[1958] emotion recognition

[1959] The device is used while the user is wearing smart glasses or a head-mounted display. The camera captures the user's facial expressions in real time, and the emotion recognition engine analyzes the user's emotions. The analysis results are sent to the server along with the selected mood and usage scenario.

[1960] Data processing and animation generation

[1961] The server analyzes the received data and generates appropriate animation data based on the selected atmosphere, usage scenario, and recognized emotion. Specifically, it combines multiple preset motions to generate smooth animation. It also dynamically adjusts the details of the animation based on the emotion information. This adjustment enables real-time animation that responds to the user's emotions.

[1962] Providing and editing animations

[1963] The animation data generated from the server is sent to the device and provided to the user on the device. The user can check the provided animation data and use editing tools to correct or adjust the animation as needed. Once editing is complete, the animation data is sent back to the server, where it is finally saved and provided again.

[1964] Specific examples

[1965] For example, when a user is smiling in a new product section during a store introduction or product introduction, the emotion recognition engine will recognize this as "joy." Based on this information, the server will generate a message saying, "Our new product is perfect for your smile!" and send it to the device.

[1966] Example prompts for generative AI models

[1967] "Generate appropriate dynamic guidance messages based on the customer's facial expression recognition data and section information. Provide the following information: emotion recognition results, the customer's current location, and summary information for each section."

[1968] As a result, this system is able to generate animations that respond to the user's emotions and provide personalized store guidance based on those animations.

[1969] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1970] Step 1:

[1971] The device activates smart glasses or a head-mounted display and captures the user's face with a camera. The input is the user's real-time facial expression data, which is captured as a camera image. This allows the device worn by the user to collect basic data for emotion recognition.

[1972] Step 2:

[1973] The emotion recognition engine installed in the device analyzes the captured facial expression data in real time and recognizes the user's emotions. The input is the camera image and the output is the user's emotion (e.g., joy, sadness, surprise, fatigue). The emotion recognition engine uses an algorithm to analyze facial features and identify emotions.

[1974] Step 3:

[1975] The device packages the recognized emotion information, along with the user-selected mood and usage scenario information, in JSON format and sends it to the server. The input is the user's emotion data and selection information, and the output is JSON data containing these. The HTTP protocol is used for transmission.

[1976] Step 4:

[1977] The server analyzes the received JSON data and generates appropriate animation data based on the selected atmosphere, usage scenario, and recognized emotion. The input is the JSON data received from the device, and the output is the generated animation data. Specifically, preset motions are combined and dynamically adjusted based on emotion information.

[1978] Step 5:

[1979] The server sends the generated animation data to the terminal. The input is the generated animation data, and the output is the animation data sent to the terminal. The HTTP protocol is used again for transmission.

[1980] Step 6:

[1981] The device presents the received animation data to the user, who then checks the animation movement. The input is animation data from the server, and the output is the animation displayed on the user's screen. This allows the user to visually check the generated animation.

[1982] Step 7:

[1983] The user can fine-tune the animation using the provided editing tools. The input is the animation displayed on the device and the user's editing operations, and the output is the animation data edited by the user. The editing tools include a function to adjust the smoothness of the motion joints.

[1984] Step 8:

[1985] The device sends the edited animation data back to the server in JSON format. The input is the animation data edited by the user, and the output is the data sent to the server. The HTTP protocol is used for transmission.

[1986] Step 9:

[1987] The server receives the edited animation data and stores it in the database. The input is the edited animation data received from the device, and the output is the data stored in the database. This way, the edited animation is saved for future use.

[1988] Step 10:

[1989] If necessary, the server sends the saved data back to the terminal so that the user can make a final check. The input is the saved animation data, and the output is the data sent to the terminal for final check. This allows the user to check the animation after the final editing is completed.

[1990] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1991] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1992] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1993] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1994] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1995] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1996] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1997] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1998] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1999] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2000] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2001] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2002] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2003] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2004] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2005] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2006] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2007] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2008] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2009] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2010] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2011] The following is further disclosed regarding the above embodiment.

[2012] (Claim 1)

[2013] A means for a user to select the mood and scene of use of the animation;

[2014] means for generating predetermined animation data based on the selected atmosphere and usage scene;

[2015] a means for providing the generated animation data to a user;

[2016] means for a user to edit the provided animation data;

[2017] a means for saving the edited animation data;

[2018] A system including:

[2019] (Claim 2)

[2020] 10. The system according to claim 1, further comprising means for combining a plurality of preset motions based on an atmosphere and a usage scene selected by a user to generate a smooth animation.

[2021] (Claim 3)

[2022] 10. The system of claim 1, further comprising means for regenerating the animation data edited by the user and providing it as saved data.

[2023] "Example 1"

[2024] (Claim 1)

[2025] A means for a user to select the mood and scene of use of the animation;

[2026] means for generating predetermined animation data based on the selected atmosphere and usage scene;

[2027] a means for providing the generated animation data to a user;

[2028] means for a user to edit the provided animation data;

[2029] a means for saving the edited animation data;

[2030] a means for utilizing a generative AI model to analyze input data and generate an optimal animation sequence;

[2031] A means of sending and receiving JSON format data between the terminal and the server,

[2032] A system including:

[2033] (Claim 2)

[2034] 10. The system according to claim 1, further comprising means for combining a plurality of preset motions based on an atmosphere and a usage scene selected by a user to generate a smooth animation.

[2035] (Claim 3)

[2036] 10. The system of claim 1, further comprising means for regenerating the animation data edited by the user and providing it as saved data.

[2037] "Application Example 1"

[2038] (Claim 1)

[2039] A means for a user to select the mood and scene of use of the animation;

[2040] means for generating predetermined animation data based on the selected atmosphere and usage scene;

[2041] a means for providing the generated animation data to a user;

[2042] means for a user to edit the provided animation data;

[2043] a means for saving the edited animation data;

[2044] means for automatically generating animations of scenes of videos and virtual reality content based on the atmosphere and usage scene selected by the user;

[2045] A system including:

[2046] (Claim 2)

[2047] 10. The system according to claim 1, further comprising means for combining a plurality of preset motions based on an atmosphere and a usage scene selected by a user to generate a smooth animation.

[2048] (Claim 3)

[2049] 10. The system of claim 1, further comprising means for regenerating the animation data edited by the user and providing it as saved data.

[2050] "Example 2: Combining Emotion Engines"

[2051] (Claim 1)

[2052] A means for a user to select the mood and scene of use of the animation;

[2053] emotion recognition means for recognizing the user's emotions in real time;

[2054] means for generating animation data based on the selected atmosphere, usage scene, and recognized emotion information;

[2055] means for providing the generated animation data to a user;

[2056] means for a user to edit the provided animation data;

[2057] a means for saving the edited animation data;

[2058] A system including:

[2059] (Claim 2)

[2060] The system of claim 1, further comprising means for combining a plurality of preset motions based on a user-selected atmosphere and usage scene, adjusting the combination based on recognized emotion information, and generating smooth animation.

[2061] (Claim 3)

[2062] 10. The system of claim 1, further comprising means for regenerating the animation data edited by the user and providing it as saved data.

[2063] "Application example 2 when combining emotion engines"

[2064] (Claim 1)

[2065] A means for a user to select the mood and scene of use of the animation;

[2066] means for generating predetermined animation data based on the selected atmosphere and usage scene;

[2067] a means for providing the generated animation data to a user;

[2068] means for a user to edit the provided animation data;

[2069] a means for saving the edited animation data;

[2070] means for recognizing a user's emotion using an emotion recognition engine and dynamically adjusting animation based on the recognized emotion;

[2071] A system including:

[2072] (Claim 2)

[2073] 10. The system according to claim 1, further comprising means for combining a plurality of preset motions based on an atmosphere and a usage scene selected by a user to generate a smooth animation.

[2074] (Claim 3)

[2075] 10. The system of claim 1, further comprising means for regenerating the animation data edited by the user and providing it as saved data. [Explanation of symbols]

[2076] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for a user to select the mood and scene of use of the animation; means for generating predetermined animation data based on the selected atmosphere and usage scene; a means for providing the generated animation data to a user; means for a user to edit the provided animation data; a means for saving the edited animation data; A system including:

2. The system according to claim 1 , further comprising means for combining a plurality of preset motions based on an atmosphere and a usage scene selected by a user to generate a smooth animation.

3. 2. The system according to claim 1, further comprising means for regenerating the animation data edited by the user and providing it as saved data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A