Information processing system

By using generative artificial intelligence models and wearable devices, personalized dance content is automatically generated and real-time feedback is provided, solving the problems of user-customized content generation and multi-user collaborative practice in existing systems, and improving the efficiency and quality of dance learning.

CN121125879APending Publication Date: 2025-12-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511213360.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-09-04
Filing Date
2025-08-28
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing dance teaching systems struggle to automatically generate customized dance content based on users' individual needs and provide real-time feedback, especially in multi-user collaborative practice scenarios where there is a lack of individual member division of labor and group coordination solutions.

Method used

By receiving user information, the system uses generative artificial intelligence models to generate personalized dance videos and music, and combines wearable devices to collect and analyze motion data in real time to provide motion correction guidance. It also automatically generates division of labor and coordination feedback in multi-user scenarios.

Benefits of technology

It enables the automatic generation and real-time feedback of personalized dance content, improving the efficiency and quality of dance learning and supporting collaborative practice in multi-user scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125879A_ABST
    Figure CN121125879A_ABST
Patent Text Reader

Abstract

The present invention provides an information processing system comprising: means for receiving information including a dance style, age, physical strength, and preference of a user; the device is used for analyzing the received information and generating original dance videos and music based on an existing dance database; and a device for transmitting the generated original dance video and music to the user terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech. Summary of the Invention

[0003] This invention provides an information processing system that receives information such as a user's dance style, age, physical strength, and preferences. The system uses a server to parse this information and, based on a dance database, generates original dance videos and music that match the user's criteria, then sends these to the user's terminal. The system further collects user movement data in real time using wearable devices, analyzes and compares it with standard movements using AI to identify problems, provides specific guidance and feedback, and guides users to correct their movements using the wearable devices. For group dances involving multiple users, the system can integrate individual inputs to generate a complete dance, assign roles to members, collect collaborative performance data in real time, analyze it, and provide precise feedback to both the group and individuals, significantly improving the efficiency and quality of dance learning and rehearsal.

[0004] "User" refers to an individual or group that uses this system to input dance information, practice, learn, or perform.

[0005] "Dance style" refers to different types or genres of dance, such as jazz dance, ballet, street dance, folk dance, etc., and is used to define the style direction of generated content.

[0006] "Age" refers to the user's age group information, which is used to determine the difficulty, range of motion, and suitability of the dance.

[0007] "Physical strength" refers to the user's physical condition, such as physical fitness, endurance, or athletic ability, which is used to determine the intensity of generated dance moves and training arrangements.

[0008] "Preferences" refer to a user's personal interests or special requirements regarding dance, music, rhythm, etc.

[0009] "Wearable devices" refer to smart devices that users wear on various parts of their bodies and that can collect motion data, such as bracelets, ankle bracelets, and headbands.

[0010] "Motion data" refers to dance-related information such as the movement, posture, and acceleration of various parts of the user's body collected through wearable devices.

[0011] "Original dance video" refers to a dance demonstration video generated by the system based on user input, which is not dependent on existing dance clips and has originality.

[0012] "Music" refers to the audio content that matches the dance moves generated by the system, including original or choreographed melodies, rhythms, and background music.

[0013] "Terminal" refers to electronic devices used by users to interact with the system, display content, and connect to external devices, including smartphones, tablets, computers, etc.

[0014] A "server" refers to a back-end processing device that receives, analyzes, performs AI calculations on, generates dance audio and video, and distributes it to terminals.

[0015] A "database" refers to a collection of data stored in a system, including dance moves, music materials, and user information, for information retrieval and processing.

[0016] "Guidance and feedback" refers to the action correction, improvement suggestions, or practice guidance information provided to users by the system based on the analysis results.

[0017] "Group dance" refers to a dance performance that involves the collaboration of multiple users and requires coordination and division of labor among members.

[0018] "Division of labor" refers to the system automatically assigning each member a specific dance move for group dances.

[0019] "Feedback" refers to the improvement suggestions, guidance information, and other content that the system returns to individuals or groups after analyzing user or group performances. Attached Figure Description

[0020] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0021] Figure 2 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0022] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0023] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0024] Figure 5This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0025] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing device and head-mounted terminal according to the third embodiment.

[0026] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0027] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0028] Figure 9 This represents an emotion map that maps multiple emotions.

[0029] Figure 10 This represents an emotion map that maps multiple emotions.

[0030] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0031] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0032] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0033] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0034] Hereinafter, an example of an implementation of the system to which the technology of this disclosure relates will be described with reference to the accompanying drawings.

[0035] First, let me explain the terminology used in the following instructions.

[0036] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0037] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0038] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0039] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0040] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0041] First Implementation Method

[0042] Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0043] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0044] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0045] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0046] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0047] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0048] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0049] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0050] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0051] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0052] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0053] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0054] Example 1

[0055] The process flow of a specific process in Example 1 will be described below. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server" and the smart device 14 is referred to as the "terminal".

[0056] Existing dance teaching and learning systems struggle to provide customized dance content and real-time feedback based on individual user characteristics (such as age, physical ability, and movement preferences). Users lack effective, automated, and intelligent movement analysis and feedback guidance when creating and practicing personalized dances. This is particularly true for multi-user collaborative practice scenarios, where systems cannot automatically generate individual member roles and group coordination plans, nor provide synchronized evaluation and improvement suggestions for individuals and teams. Therefore, how to automatically generate dance content based on personalized needs, combined with real-time analysis of movement data and individualized, team-based feedback guidance, is a pressing technical problem that needs to be solved in this field.

[0057] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0058] In this invention, the server includes: an information input unit for inputting motion information, age information, physical ability information, and preference information from an information input device; a data conversion and transmission unit for converting user information into a predetermined data format and transmitting it through a communication network; a generation unit for retrieving motion pattern information and audio information from an information storage device based on user information, and generating personalized motion images and audio information using a generative artificial intelligence model combined with prompts; an information providing unit for sending personalized motion images and audio information to a terminal device; and an output unit for displaying and playing the aforementioned information by the terminal device. The server may also include: a data acquisition unit for acquiring real-time motion data from a motion monitoring device, and an analysis, generation, and assistance unit for analyzing and comparing the data with generated standard motion images and generating feedback; it may also include an information integration, analysis, and allocation unit for automatically generating integrated dances and allocating individual motion information for multi-user scenarios based on group identification information, and generating evaluations and improvement suggestions for teams and individuals.

[0059] This allows for the automatic generation of original dances and background music based on users' personalized needs, enabling efficient and intelligent dance content recommendations, movement analysis, and correction guidance for individuals and teams. This improves the relevance and efficiency of dance learning while effectively supporting collaborative practice feedback in multi-user scenarios.

[0060] "Information input device" refers to hardware or software interface used to receive and input personalized information such as user action information, age information, physical ability information, and preference information.

[0061] "Motion information" refers to all data that can characterize the characteristics of a user's dance or sports movements, such as descriptions, categories, and parameters.

[0062] "Age information" refers to data used to represent a user's age or age group.

[0063] "Physical ability information" refers to data that characterizes a user's physical condition, such as stamina, flexibility, and strength.

[0064] "Preference information" refers to data on users' personal interests, choices, or specific needs in areas such as action type, dance style, and music genre.

[0065] The “data conversion and transmission unit” refers to the functional component that converts input information into a standard format and transmits it to the server via a network.

[0066] "Information storage device" refers to hardware equipment or database system used to store various necessary information such as action patterns, audio, and historical data.

[0067] "Generative AI models" refer to AI algorithms or systems that can automatically generate personalized content (such as motion graphics, music, etc.) based on input prompts.

[0068] "Prompt statements" refer to text or data information used as input conditions for generative artificial intelligence models to limit or describe the content to be generated.

[0069] "Personalized motion graphics" refers to dance videos, animations, or motion image content that are customized and generated based on user needs and motion data.

[0070] "Audio information" refers to the audio content such as background music and rhythmic cues provided in conjunction with the action visuals.

[0071] "Information providing unit" refers to a hardware module or software program used to push generated content from the server to the terminal device.

[0072] "Terminal device" refers to the hardware device used by users to receive, display, and play dance content, such as smartphones, tablets, and personal computers.

[0073] "Output unit" refers to a device component that can display the received image and audio data to the user in a visual and audible manner.

[0074] "Motion monitoring information acquisition device" refers to a sensing device worn on the user's body that can collect human motion data in real time, such as smart bracelets and motion capture equipment.

[0075] "Data acquisition unit" refers to a component or module used to receive and collect motion data from motion monitoring devices.

[0076] The "information analysis unit" refers to a functional module used to process and compare the user's actual actions with standard actions, and to analyze the differences and problems in the actions.

[0077] The "information generation unit" refers to the functional module that generates improvement suggestions, guidance content, and feedback data based on the analysis results.

[0078] "Assistive unit" refers to a device or module used to communicate with and guide motion adjustments to the user through visual, auditory, or tactile means.

[0079] "Group identification information" refers to data used to distinguish and identify multiple users belonging to the same practice or performance team.

[0080] The "Information Integration and Distribution Unit" refers to the functional module that enables information integration, division of labor, and content distribution in multi-user scenarios.

[0081] The “Analysis and Evaluation Unit” refers to a functional module used to analyze and evaluate team collaboration performance and the accuracy of individual actions.

[0082] "Information Providing Unit (Team Scenario)" refers to a module that generates and pushes improvement suggestions, feedback data, and other content to teams or individuals in team mode.

[0083] To better understand and implement this invention, the following provides a detailed description of its embodiments.

[0084] The system described in this invention mainly includes hardware such as a server, terminal, information input device, motion monitoring information acquisition device (such as a wearable device), and information storage device, as well as software modules such as a generative artificial intelligence model, a database management system, and terminal application software. Data interaction between the various units is achieved via a communication network. Its main functions are: to automatically generate personalized dance content for users, to intelligently analyze motion tracking data, and to output feedback for individual and team improvements.

[0085] In this system, users operate the terminal application software through information input devices (such as smartphones, tablets, and personal computers) to input information such as their movement needs, age, physical abilities, and dance and music preferences. This input interface can be a chat window, a form, or voice interaction. The terminal application formats the user information into standard data, such as JSON, and sends it to the server via the network. Specific examples of information input devices include general-purpose smart terminal devices such as Android smartphones, iOS smartphones, general-purpose tablets, and personal computers.

[0086] After receiving user input data, the server parses and cleans the data to obtain clearly defined structured parameters. The server then reads dance template libraries, dance movement databases, and music databases from its information storage device (e.g., relational databases such as MySQL or PostgreSQL). Subsequently, the server invokes generative artificial intelligence models, such as deep learning-based motion generation networks, language-action translation networks, and natural language processing modules (software tools such as PyTorch and TensorFlow), inputting the user-supplied parameters and prompts into the AI ​​model. Based on user needs, the server automatically generates personalized dance movement videos and original audio. For example, the server can utilize audio and video media processing tools such as FFmpeg to synthesize and post-process video and audio sequences.

[0087] The generated personalized dance videos and music files (such as mp4 and mp3) are pushed to the user's terminal via the network. After receiving them, the application on the terminal calls multimedia controls (such as MediaPlayer and AVPlayer components) to seamlessly present the dance content and provides users with replay, pause, slow-motion, and other operations, as well as text or voice instructions for each step of the movement. The terminal also supports starting the practice module at any time during the demonstration.

[0088] When practicing dance, users need to wear wearable devices for motion monitoring (such as smart bracelets, smart ankle bracelets, body motion capture straps, etc.) as prompted by the system. After the user selects the "Practice" or "Real-time Tracking" function, the terminal communicates with the wearable device via Bluetooth or Wi-Fi to obtain raw data such as three-axis acceleration and angular velocity of key parts of the user's body. At the same time, the camera can record video of the dance process. The terminal performs basic data processing and then uploads it to the server.

[0089] The server receives motion and video data, invokes motion recognition and analysis algorithms (such as OpenPose and MediaPipe skeletal point detection algorithms), and automatically matches and compares the user's actual movements with the generated standard dance movements. The server integrates real-time data to analyze differences in user movements, identifying issues such as movement deviations and rhythm synchronization, and automatically generates detailed "movement improvement suggestions" or "personalized feedback" using a natural language processing model. The server can send this feedback to the terminal in text, voice, or multimedia format, which displays it on the screen and, as needed, controls wearable devices to provide haptic cues (such as vibration and indicator lights) to help the user correct their movements promptly.

[0090] When multiple users practice or perform as a team, the system allows users to input a group identifier code and their individual characteristic parameters. After collecting all member information, the server collaboratively calls the AI ​​model to automatically generate a team dance plan with collaborative division of labor. Each member will receive their assigned movements and audio. Based on each user's actual training data and overall coordination data, the server outputs team improvement suggestions and individual movement evaluations, which are pushed to each member's terminal. The terminal application can display individual and group progress curves, performance summaries, and specific improvement suggestions.

[0091] Specific hardware examples:

[0092] The server can be a host or cloud server based on a high-performance CPU and GPU.

[0093] The terminal device can be any type of general-purpose mobile terminal equipment.

[0094] Motion capture devices can include smart bracelets with IMUs, customized motion sensors, motion capture systems, etc.

[0095] Specific software examples:

[0096] Generative artificial intelligence models can be used for action generation networks, GPT-like natural language generation and analysis models, and deep learning-based 3D action modeling tools, among others.

[0097] Database software can be a relational database.

[0098] Video and audio synthesis tools can be found in multimedia processing software libraries.

[0099] The terminal application can be a multi-platform app.

[0100] Specific examples of usage:

[0101] A user entered the following into the app: "Design a fast-paced original jazz dance routine for users aged 20 with intermediate skill level, accompanied by original music."

[0102] For team activities, users can enter in the app: "Please design a lively and energetic dance suitable for a wedding for a team of 5 people, with each person having a different dance role."

[0103] Through the aforementioned comprehensive hardware and software integration, automatic dance generation, personalized movement analysis, instant feedback, and real-time guidance for multi-user collaboration are achieved, thereby significantly improving user experience and dance learning efficiency.

[0104] use Figure 11 The processing procedure is explained.

[0105] Step 1:

[0106] Users open the dance customization application on their devices, click "Create New Dance Customization," and enter information such as dance style, age, physical ability, and music preferences.

[0107] Input: Text related to user needs, such as "20 years old, likes fast-paced jazz dance, intermediate level".

[0108] After receiving the input, the terminal preprocesses the data, formats it into JSON, and prepares it for uploading.

[0109] Output: Formatted user-customized requirements data.

[0110] Step 2:

[0111] The terminal transmits formatted data to the server via an encrypted communication network.

[0112] Input: Formatted custom requirement data.

[0113] The terminal confirms the network status and securely sends the data using protocols such as HTTPS.

[0114] Output: Notification of successful data upload; data has arrived at the server.

[0115] Step 3:

[0116] The server parses the user-customized request data uploaded by the terminal and retrieves relevant action templates and audio materials from the information storage device (database).

[0117] Input: User-required JSON data.

[0118] The server parses the fields, mapping parameters such as dance type, speed, and difficulty to database search criteria to retrieve relevant information.

[0119] Output: Initial selection of action templates and audio materials that match the requirements.

[0120] Step 4:

[0121] The server invokes a generative artificial intelligence model, taking user requirements, initial selected materials, and prompts as input, to automatically generate original dance movement data and audio files.

[0122] Input: User requirements parameters, initial materials selected from the database, and prompts from the generative AI model (e.g., "Please generate fast jazz dance moves and background music suitable for a 20-year-old woman").

[0123] The server converts natural language parameters into model input, and the AI ​​model processes them to generate personalized motion videos and audio.

[0124] Output: Generated MP4 dance video, MP3 background music, and motion description document.

[0125] Step 5:

[0126] The server will generate dance videos, audio recordings, and movement instructions, which will then be transmitted to the user's terminal via the network.

[0127] Input: Multimedia files and descriptions output by the AI ​​model.

[0128] The server organizes the output content and pushes it to the terminal application in sequence, monitoring the transmission success rate.

[0129] Output: The terminal receives dance video, audio, and motion text.

[0130] Step 6:

[0131] The terminal receives and displays dance videos and audio, allowing users to begin learning and practicing.

[0132] Input: dance video file, audio file, and motion prompt text.

[0133] The terminal uses a multimedia playback component to display content, showing text descriptions and supporting operations such as slow motion, pause, and replay. The user prepares according to the terminal prompts and clicks "Start Practice".

[0134] Output: The user plays the video and enters the practice preparation state.

[0135] Step 7:

[0136] Users can wear motion monitoring wearable devices, such as smart bracelets or motion capture belts, according to the terminal prompts, and start motion training.

[0137] Input: Wearing complete notification, user clicks "Start Practice".

[0138] Users can mimic the dance moves in the video.

[0139] Output: Wearable device data acquisition is enabled, and real-time motion data stream is generated.

[0140] Step 8:

[0141] The terminal receives motion sensing data collected by wearable devices in real time and records video synchronously using a camera.

[0142] Inputs: IMU sensor data stream, camera video recording data.

[0143] The terminal performs data integration (such as time synchronization and format sorting) and uploads it to the server periodically.

[0144] Output: The aggregated motion sensor data and video clips are uploaded to the server.

[0145] Step 9:

[0146] The server receives user motion data and practice videos, and uses motion recognition and analysis algorithms to compare them frame by frame with standard motions generated by AI to analyze motion deviations and rhythm differences.

[0147] Input: User's actual movement data, video recordings, and standard dance movement data.

[0148] The server performs processes such as skeletal point extraction, time sequence analysis, angle change statistics, and rhythm alignment to identify problem areas.

[0149] Output: Detailed action difference report and personalized improvement suggestion text.

[0150] Step 10:

[0151] The server will send improvement suggestions and feedback information to the terminal via the network.

[0152] Input: Results and suggestions of motion difference analysis.

[0153] The server pushes feedback content such as text, voice, and images to the terminal app.

[0154] Output: Suggestions for improving the terminal display and real-time feedback.

[0155] Step 11:

[0156] The terminal displays analysis results and improvement suggestions to users, and guides and corrects movements through vibrations and other means using wearable devices.

[0157] Input: Feedback text, voice, or haptic prompts.

[0158] When the terminal displays a red warning, broadcasts an audio message, or the wearable device vibrates briefly to alert the user when the device is in an incorrect position, the device will remind the user.

[0159] Output: Users receive specific improvement suggestions and real-time correction assistance.

[0160] Step 12:

[0161] In the case of a multi-person team mode, users enter their individual needs and bind them to the same group ID, and the terminal merges and uploads all team member information to the server.

[0162] Input: Team member information, group ID, and individual customized parameters.

[0163] The terminal integrates the relevant data and uploads it as a whole.

[0164] Output: The server received complete team information.

[0165] Step 13:

[0166] The server automatically generates the overall dance choreography and the division of labor for each member based on team information. Then, it synchronously analyzes the coordination and cooperation based on real-time motion data, and generates improvement suggestions for both the group and individuals.

[0167] Input: Team requirement parameters, multi-member actions and status data.

[0168] The server allows the AI ​​model to generate group dance steps and individual tasks that fit the team performance, and outputs comprehensive comments and improvement suggestions based on the team's coordination.

[0169] Output: Team dance division video, individual movement tips, and team improvement suggestions are pushed to each member's terminal.

[0170] Step 14:

[0171] The terminal displays collective and individual suggestions to each team member and controls local wearable devices to provide feedback based on individual needs, helping the team coordinate practice efficiently.

[0172] Input: Team and individual feedback information pushed by the server.

[0173] The terminal displays team performance charts and individual evaluations simultaneously, and provides more frequent device alerts for members prone to errors.

[0174] Output: Team members receive feedback simultaneously, guiding them to complete group exercises efficiently.

[0175] Application Example 1

[0176] The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0177] In existing technologies, it is difficult for individuals or groups to personalize and optimize movement patterns based on specific environments, individual differences, and emotional states when engaging in physical movement training, dance instruction, or operating automated equipment. The lack of efficient real-time data acquisition and feedback mechanisms leads to insufficient efficiency and adaptability in movement training or automation processes. Users struggle to obtain personalized improvement suggestions, and bottlenecks exist in equipment collaboration or group movement coordination. Therefore, there is an urgent need for an information processing system that can combine the real-time status of users or equipment, generate optimal movement plans through artificial intelligence, and provide dynamic feedback to improve the intelligence and personalization of overall movement guidance and automated operation.

[0178] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0179] In this invention, the server includes: a device for acquiring environmental information containing user body movement information and preference conditions via a terminal; a device for parsing the environmental information and inputting prompts into a generative artificial intelligence model to generate optimal movement patterns and related content; a device for sending the generated information and movement patterns to a terminal or device; a device for real-time acquisition and collection of user or device movement data and uploading it to the server; a device for analyzing the real-time acquired data and generating evaluation and feedback information; and a device for sending the feedback information to the terminal or device to guide correction and improvement. This enables personalized automatic generation of movement patterns and content, combined with real-time data and emotion analysis for intelligent feedback, effectively improving the user's or device's operational efficiency, movement accuracy, and the level of intelligence and personalization of the user experience.

[0180] "Terminal" refers to a data processing device that can interact with users and be used to input or output information, including but not limited to smartphones, tablets, and personal computers.

[0181] "User" refers to a human individual or participant who uses this system for motion training, data input, or to receive feedback.

[0182] "Environmental information" refers to data related to the current physical, emotional, or operational conditions of a user or device, including temperature, humidity, spatial layout, and operational requirements.

[0183] "Body motion information" refers to motion-related data such as the type of movement, range of motion, and rhythm of movement exhibited by a user or device during an activity.

[0184] "Preference conditions" refer to users' interests and personalized needs in terms of movements, dance, music, style, rhythm, etc.

[0185] "Generative artificial intelligence model" refers to an artificial intelligence algorithm or system that can automatically generate action patterns, content or sound information that meet the requirements based on input prompts.

[0186] "Prompt information" refers to data or statements used as input to generative artificial intelligence models to clarify the task type, requirements, or constraints.

[0187] "Action pattern" refers to a series of continuous action processes or action plans that are pre-optimized and generated for a specific target.

[0188] "Audio information" refers to the audio data such as music, rhythm cues, and sound effects provided in accordance with the action mode or instruction content.

[0189] "Output device" refers to a device used to display system-generated content to users, including displays, speakers, wearable devices, etc.

[0190] A "server" refers to a central information processing system used for the entire process of data reception, storage, analysis, generation, and feedback.

[0191] "Actual motion data" refers to the status and motion-related data that a user or device acquires in real time through sensing devices during the actual execution of actions.

[0192] "Feedback information" refers to suggestions or instructions output by the system after analyzing and comparing the user's actual actions with the target actions, which are used to guide the user to adjust and improve their operations.

[0193] "Wearable body information collection devices" refer to devices worn by users on different parts of the body that can be used to collect physiological and motion signals in real time, such as sensor bracelets, smartwatches, motion capture devices, etc.

[0194] "Emotional state" refers to the psychological reaction or emotional type identified by analyzing a user's physiological signals, facial expressions, or behavior, such as pleasure, tension, and anxiety.

[0195] "Personalized optimal information" refers to action, instruction, or audio content generated and presented in the most suitable way for an individual based on their personal conditions.

[0196] "Group action data" refers to a collection of action-related data collected from multiple users that reflects the overall collaborative effect.

[0197] "Evaluation" refers to the system's objective judgment of the comparison between user or equipment actions and performance and target reference standards.

[0198] "Guidance" refers to the suggestions for action adjustment, practice methods, or improvement that the system outputs to the user based on the evaluation results.

[0199] This invention provides a motion generation and intelligent feedback system based on a generative artificial intelligence model and prompts, which can be widely applied in various fields such as individual physical training, dance instruction, and collaborative robot motion optimization. The system includes a server, a terminal, a user, wearable devices, and multiple data acquisition, processing, and feedback modules. The following description, in conjunction with the main hardware and software and actual usage, supports the scope of the claims of this invention.

[0200] The system for implementing the present invention may be composed of the following devices:

[0201] Users run specific front-end applications using terminal devices (including smartphones, tablets, personal computers, etc.). These applications are based on web front-end frameworks (such as Vue.js and React) or mobile application development environments (such as Android Studio and Xcode). The terminal is responsible for receiving environmental information, body movement information, and personal preferences input by the user. After filling in detailed data such as dance style, movement goals, age, physical strength, interests, and real-time mood through the interface, the user clicks the send button.

[0202] The terminal formats and processes user input information, converting it into a standardized data structure, and then sends it to the server via a wireless or wired network. The server can be deployed in data centers, cloud platforms, or edge computing nodes, and typical server hardware includes high-performance processors, multi-core memory, and efficient network interfaces.

[0203] The server-side has several built-in software modules: a data receiving and parsing module (based on Python or Java), a generative artificial intelligence model module (using deep learning frameworks such as Keras and TensorFlow), an emotion recognition module (which can use deep learning-based sentiment analysis tools), a database management module (such as MySQL or MongoDB), and a data output and feedback distribution module.

[0204] After parsing the received data, the server inputs specific prompts and parameters into the generative artificial intelligence model to achieve the following functions: automatically generating action plans or dance content that best match the environment and individual characteristics; generating sound information (such as music and sound effects) that are adapted to the actions or training content; and subdividing the generation into various modes such as individual training and collaborative group tasks according to needs.

[0205] The generated content is distributed back to the user terminal via API interfaces and security protocols. The terminal displays original dance videos, motion instruction text, or audio content to the user in a multimedia format. For robotic task scenarios, the server can output the optimal control sequence to the operating device via an industrial bus protocol.

[0206] During system implementation, users can wear wearable devices (such as smartwatches, sensor bracelets, and motion capture devices) on multiple parts of their bodies as needed. These devices communicate with the terminal in real time via Bluetooth, Wi-Fi, and other communication methods. These devices can collect the user's body's motion data and physiological signals, such as wrist and ankle acceleration and heart rate.

[0207] The terminal uploads the collected data (including video streams of user actions captured by the camera) to the server. The server uses an artificial intelligence model to identify and analyze the user's current state and actions, compares the actual actions with the generated content, analyzes differences in actions and potential problems, and, in conjunction with the emotion recognition module, produces customized feedback information.

[0208] Feedback generated by the server will be sent back to the terminal or device via the network. The terminal can remind the user of actions that need improvement or convey emotional encouragement through various means such as text, voice, images, and vibration. In robot control scenarios, the system can automatically optimize motion parameters to achieve adaptive operation and efficiency improvement of factory equipment.

[0209] When multiple users participate, the system supports collecting individual data and group goals through the terminal. The server generates collaborative action content and distributes individual guidance suggestions based on the status of all members, achieving optimal actions and efficiency improvement in team collaboration.

[0210] In practical use, users only need to fill in their needs and status through the terminal interface, wear the wearable device, and follow the terminal's guidance to train or operate. They can then repeatedly receive intelligent guidance based on the latest actions and emotional states, as well as personalized action generation and feedback. The system features multiple innovative characteristics, including real-time analysis, adaptive generation, dual optimization of groups and individuals, and integration with emotions.

[0211] In practical application scenarios, users can input natural language prompts such as "The factory temperature is 22℃, the humidity is 40%, the layout is Type A, the robot is a robotic arm type, medium speed, high precision" or "Jazz dance style, pleasant mood, intermediate level, suitable for young people in their 20s, who like fast-paced music" through the terminal, and the system can automatically complete the entire process of action or content generation and feedback.

[0212] use Figure 12 The processing procedure is explained.

[0213] Step 1:

[0214] Users open the application via a terminal (such as a smartphone, tablet, or computer), and input environmental information, body movement information, and preferences on the interface, such as dance style, age, physical strength, interests, current mood, or temperature, humidity, layout, and robot type in a factory setting. The user clicks the "Send" button. Input consists of various structured or natural language text data. Output is the raw information data input by the user.

[0215] Step 2:

[0216] After receiving user input, the terminal formats the input, converting text and parameters into standardized data structures (such as data objects mapped to specified fields), and performs preliminary validation. The input consists of various types of information entered by the user. The terminal encapsulates this data in formats such as JSON and then sends the data packets to the server over the network. The output is a structured data packet.

[0217] Step 3:

[0218] After receiving the data packet from the terminal, the server first uses the data parsing module to read and verify all fields to ensure the data is complete and error-free. The input is the structured data packet uploaded by the terminal. The server combines the extracted parameters into a prompt statement for subsequent data processing and modeling. The output is the prepared prompt message and feature parameters.

[0219] Step 4:

[0220] The server inputs prompts and relevant parameters into a generative AI model (such as a deep learning network built on TensorFlow or Keras), performing preprocessing operations such as feature vector encoding and data normalization on the input. The input consists of user environment information and individual parameters. Based on the input data, the generative AI model automatically infers and outputs optimal action patterns, original action content (such as action sequences or dance videos), or audio information.

[0221] Step 5:

[0222] The server processes the content generated by the model, producing files in various formats suitable for terminal playback, such as videos and motion instructions interpretable by mechanical equipment. The server then sends these files or data to the terminal or robot control unit via the network. The input is the content generated by the AI ​​model. The output is adapted multimedia data or motion instructions.

[0223] Step 6:

[0224] After receiving action content, video, or audio information from the server, the terminal automatically plays the video on the interface, displays action guidance text, or provides an interactive operation entry point for the user. In a robot factory scenario, the robot actually performs operations according to the sequence of instructions issued by the server. The input is the generated content returned by the server. The output is a multimedia display that is visual, audible, or interactive for the user, or the actual actions performed by the robot.

[0225] Step 7:

[0226] When users practice movements, dance, or perform operations, they wear wearable devices (such as smart bracelets, motion capture devices, etc.). The terminal collects the user's motion data, heart rate, facial expressions, or the sensor status of the device in real time. The input is the user's own physiological and movement signals. The terminal integrates, encapsulates, and uploads the collected data to the server. The output is a real-time behavior data packet.

[0227] Step 8:

[0228] The server utilizes a motion data analysis module and a sentiment analysis engine to receive and parse real-time data from the terminal. It compares the actual performance with specified target movements in the database to identify deviations between movement and emotion. The input is real-time collected motion and emotion data. Based on this, the server performs data analysis and model comparison, outputting personalized, real-time correction suggestions and emotional encouragement feedback.

[0229] Step 9:

[0230] The server sends the generated feedback information to the user terminal or robotic device via the network. The input is the feedback information analyzed by the server. After receiving the feedback, the user terminal displays correction suggestions on the interface, such as specific action improvement prompts and encouragement messages; wearable devices can also provide vibration or light signal prompts. The output is user-perceptible and actionable feedback content, guiding the user to further adjust actions or parameters and enter the next round of practice and optimization.

[0231] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0232] Example 2

[0233] The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0234] Existing dance generation and training support systems struggle to fully reflect users' personalized preferences and emotional states, and cannot adjust movement corrections and feedback during practice in real time based on user emotions. Furthermore, for multi-person dance generation and instruction, current technologies are inadequate in terms of movement allocation, team coordination analysis, and diverse feedback for individuals and the group as a whole. There is a lack of comprehensive solutions that can combine sentiment analysis and generative artificial intelligence models for intelligent content generation, precise analysis, and adaptive feedback.

[0235] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0236] In this invention, the server includes devices for acquiring information on the user's movement type, age group, physical ability, preferences, and emotional state; devices for generating movement video and audio data adapted to the user's needs based on an information storage device and a generative artificial intelligence model; and devices for sending the generated content to an information display device. The server also includes devices for acquiring and parsing temporal movement information collected by a wearable input device and comparing it with standard movement data to automatically identify problematic parts of the movement; and devices for generating adaptive guidance and feedback based on the user's emotional state identified by an emotion analysis device. Simultaneously, the server can integrate information input by multiple users, automatically allocate collective and individual movement content, and generate guidance and feedback for teams and members respectively through movement and emotion evaluation processing. This enables highly customized dance generation and intelligent practice guidance for individuals or teams, and provides accurate and effective movement correction and positive incentive feedback in real time based on the user's actual movements and emotional state, improving the dance training experience and learning efficiency.

[0237] "Motion type" refers to the dance or sports category selected or specified by the user, including but not limited to dance style, movement style, etc.

[0238] "Age range" refers to the age range of the user, which is used to determine the appropriate difficulty of the movements and the performance content.

[0239] "Physical ability" refers to the user's physical fitness level, such as endurance, flexibility, and coordination, which is used to adjust the intensity and complexity of movements.

[0240] "Preference" refers to a user's personalized inclination or liking for dance, music, rhythm, etc.

[0241] "Emotional state" refers to the emotional state that the user is currently experiencing during performance or practice, including feelings of pleasure, tension, and drowsiness.

[0242] "Wearable input devices" refer to electronic devices that can be worn by users and used to collect motion data in real time, including but not limited to wristbands, smartwatches, etc.

[0243] "Time-sequence motion information" refers to motion sensor data collected sequentially over time, used to reconstruct the user's dynamic motion process.

[0244] "Information processing device" refers to an electronic computer system that carries out data analysis, content generation and control functions.

[0245] "Information storage device" refers to a data storage medium used to store content data such as dance and music, as well as user-related information.

[0246] "Information display device" refers to terminal equipment that displays generated content and feedback information, including smartphones, tablets, computers, etc.

[0247] "Generative artificial intelligence models" refer to artificial intelligence systems and algorithms that automatically generate action sequences or audio content based on input data.

[0248] "Cue data" refers to descriptive information input into generative artificial intelligence models, including action needs, emotional needs, and other content.

[0249] "Motion video data" refers to digital video files generated by the system to showcase specific dances or movements.

[0250] "Audio data" refers to the accompanying music or related sound files generated by the system according to user needs.

[0251] "Control signals" refer to electronic signals used to adjust the status of equipment or guide user actions, including vibration and light signals.

[0252] "Motion integration processing" refers to the process of synthesizing information input from multiple users into team performance data and formations.

[0253] "Action composition information" refers to the unique action descriptions and content assigned to different users.

[0254] "Motion evaluation processing" refers to the method of analyzing collected motion data and judging its completion quality and teamwork level.

[0255] "Emotion analysis device" refers to a device or module used to analyze and identify a user's emotional state.

[0256] "Guidance content and response information" refers to messages such as action correction suggestions and emotional feedback generated by the system, which are used to guide users to practice and improve their performance.

[0257] To facilitate others' understanding and implementation of this invention, the embodiments of this invention are described in detail below.

[0258] This system comprises servers, terminals, and users, and can be composed and executed through various hardware and software components. Servers are typically deployed on high-performance computing devices, such as general-purpose servers or cloud platforms, integrating data storage devices. Terminals can be smartphones, tablets, or personal computers, and information display devices. Users interact with the system through terminals. Wearable input devices include wristbands, smartwatches, and motion sensors. The system relies on multiple software modules working together, including modules for information acquisition, data parsing, generative artificial intelligence model inference, sentiment analysis, content generation, and feedback. Commonly used software and tools include Linux server operating systems, artificial intelligence inference frameworks (such as PyTorch and TensorFlow), database systems (such as MySQL and PostgreSQL), and front-end mobile applications.

[0259] The server in this invention first acquires user input information, including action type (such as dance style), age group, physical ability, preferences, and emotional state. The terminal transmits the user input data to the server via a standard interface. After parsing this information, the server uses an information storage device (database) to invoke a generative artificial intelligence model. This model can be a text-to-action dance generation model, a text-to-music generation model, or a sentiment analysis model; specific examples include a Transformer-based text generation model, a diffusion-based image / video generator, and an audio synthesis system. The server inputs user request information as prompts to the corresponding generative artificial intelligence model.

[0260] For example, a user inputs the following prompt on the terminal: "Generate an original dance breakdown video and original soundtrack that incorporates jazz dance elements, a cheerful mood, is suitable for women in their 20s, has a moderate difficulty level, and is accompanied by fast-paced electronic music." The server combines this text with relevant dance knowledge and movement templates in its database to automatically generate dance video data and corresponding music and audio data. The server then uses motion recognition models such as OpenPose or MediaPipe to further analyze and optimize the generated motion sequence, ensuring that the movements are smooth and ergonomic throughout.

[0261] The server is also responsible for uploading the generated content to the information display device. After receiving the dance video and music, the terminal displays it, allowing users to practice step-by-step at any time. To enhance interactivity, the system supports wearable input devices to collect the user's dynamic motion data (such as acceleration, angular velocity, and spatial positioning) in real time. The terminal collects this temporal motion information and uploads it synchronously to the server.

[0262] The server analyzes this real-time motion data, compares it with generated standard motion data, and uses pattern recognition or skeletal motion modeling software to automatically identify differences and problem areas. For example, it detects deviations between the user's movements and standard movements in terms of footwork, gestures, and rhythm. The server combines the motion parameters collected by the wearable device with the judgment results from emotion analysis devices (such as facial expression and emotion analysis modules) to generate customized feedback suggestions and encouraging language for specific problems and emotional states. The server further sends control signals to the wearable input device, such as guiding the user to correct specific limb movements through vibration.

[0263] In multi-user scenarios, the server can simultaneously collect personalized information from multiple users, automatically integrate and calculate the team's overall performance movements and music, and assign different movement components to each member. After collecting the movement data of each member, the server comprehensively evaluates the synchronization of teamwork, overall performance, and accuracy of individual movements, and then automatically generates collective and individual guidance information and feedback suggestions based on the results, improving the efficiency and engagement of team dance training.

[0264] For example, multiple users might input the following prompt: "Please choreograph a dance for a four-person group, focusing on a joyful mood and incorporating elements of jazz and modern dance. A prefers high-difficulty spins, B favors elegant movements, and the other two should generate segmented movements and background music based on general preferences." The server will then generate a video and music for the entire team's choreography, automatically assign movement segments to each member, and provide personalized follow-up suggestions and positive feedback.

[0265] In summary, this invention combines multiple types of hardware devices and high-performance software modules to dynamically expand user needs into original dances and background music adapted to different physical and emotional states. Through analysis and feedback empowered by artificial intelligence technology, it achieves a highly personalized, interactive, and intelligent dance generation and training support system.

[0266] use Figure 13 The processing procedure is explained.

[0267] Step 1:

[0268] Users launch the application via a terminal (such as a smartphone, tablet, or personal computer) and input information including movement type, age group, physical ability, preferences, and emotional state. The input can be a text description, such as "Generate a cheerful style jazz dance suitable for a 20-year-old woman to fast-paced music."

[0269] Input: Text or options related to the user's requirements.

[0270] Output: Structured user requirement data.

[0271] Specific actions: The user fills in the options, enters a description, and clicks the "Submit" button.

[0272] Step 2:

[0273] The terminal receives user input, formats the data (e.g., converts it into a unified JSON or other standardized data structure), and performs preliminary verification of the content integrity.

[0274] Input: The original input information submitted by the user.

[0275] Output: Formatted, fully structured data.

[0276] Specific actions: The terminal performs data format checks, prompts for non-compliant content, and calls the local formatting API to encapsulate the data.

[0277] Step 3:

[0278] The terminal securely uploads the formatted user request data to the server via the network interface.

[0279] Input: Formatted user requirement data.

[0280] Output: Data packets submitted to the server and upload results.

[0281] Specific actions: The terminal calls the network transmission module to send a request via HTTP / HTTPS protocol, and the server confirms receipt and returns a status.

[0282] Step 4:

[0283] The server parses the received data, requests the internal information storage device (database) to retrieve basic actions, music templates and content related to the requirements, and constructs prompt statements based on the retrieved results for subsequent generation.

[0284] Input: A formatted user requirement data package.

[0285] Output: Organized prompts and retrieved relevant template information.

[0286] Specific actions: The server calls a database query function to concatenate the input parameter text of the generative artificial intelligence model.

[0287] Step 5:

[0288] The server calls a generative artificial intelligence model, inputs the prompts into the model, performs content generation inference, and outputs original motion sequences (such as skeletal point sequences), dance videos, or customized audio.

[0289] Input: Prompt statement and template data.

[0290] Output: Generated dance video and audio data.

[0291] Specific actions: The server initializes the AI ​​model inference environment, inputs parameters, waits for the model inference to finish and obtains the output.

[0292] Step 6:

[0293] The server further analyzes and corrects the AI-generated content, uses motion recognition or posture analysis modules to enhance motion smoothness, and adds tags to optimize the content based on the user's emotional needs.

[0294] Input: The generated video and audio data.

[0295] Output: Optimized motion video and audio data with emotional tags.

[0296] Specific actions: The server calls motion analysis software to correct, score, and label the AI-generated results frame by frame.

[0297] Step 7:

[0298] The server pushes the generated and optimized motion video and audio data to the user's terminal through a content distribution mechanism.

[0299] Input: Final dance video and audio data.

[0300] Output: Data packet is uploaded to the terminal and the reception status is confirmed.

[0301] Specific actions: The server uploads video and audio files to the cloud or local ID storage and pushes download links to the terminal.

[0302] Step 8:

[0303] After receiving the dance video and music, the terminal calls the local multimedia player module to display the content, providing users with step-by-step learning and full-process playback functions.

[0304] Input: The video and audio content returned by the server.

[0305] Output: The terminal interface displays dance videos and music playback, which users can watch and learn from.

[0306] Specific actions: The terminal loads multimedia data, decodes it, and displays the playback interface, supporting operations such as pause, fast forward, and slow motion.

[0307] Step 9:

[0308] After wearing the wearable input device, the user can start the motion training mode through the terminal, synchronously collect real-time motion data (such as acceleration, angular velocity, spatial positioning, etc.) and record it, and upload it to the server via Bluetooth or wireless protocol.

[0309] Input: The user's actual movements and video data during the movement process.

[0310] Output: Real-time captured and uploaded sequence action and video data streams.

[0311] Specific actions: When the user clicks "Start Practice", the terminal starts sensor data acquisition and camera recording, and multiple signals are fused, packaged, and uploaded.

[0312] Step 10:

[0313] The server receives action and video data, performs multi-dimensional feature comparison between the user's real-time performance and standard reference actions, and combines the emotion analysis module to detect the user's facial expressions and emotions, outputting action errors, emotion deviations and improvement suggestions.

[0314] Input: User actions and video data stream.

[0315] Output: Action problem analysis results, sentiment analysis results, and feedback suggestions.

[0316] Specific actions: The server performs data preprocessing, calls the AI ​​pose comparison and expression recognition API to output quantified error and text description.

[0317] Step 11:

[0318] The server pushes the analysis results and feedback suggestions to the terminal in the form of text, instructions or warning signals. When necessary, it also sends vibration control signals to the wearable device to guide the user to make action corrections and emotional adjustments in real time.

[0319] Input: Feedback and signals generated by the server.

[0320] Output: Terminal interface feedback, physical reminders from wearable devices.

[0321] Specific actions: The terminal displays personalized learning suggestions and vibration alerts; the user makes adjustments based on the alerts.

[0322] Step 12:

[0323] In multi-user scenarios, the server receives input from multiple users, analyzes it comprehensively, automatically generates the overall dance content for the team, assigns individual movements, collects the performance of each member, analyzes the degree of collaboration and synchronization, and then outputs diverse feedback and suggestions to each member or the team as a whole.

[0324] Input: Multi-user demand data and training data stream.

[0325] Outputs: Team action assignments, individual and group performance evaluations and feedback.

[0326] Specific actions: The server performs grouping and calculations based on a preset algorithm, generates content for each group, and automatically pushes it to the relevant terminals.

[0327] Application Example 2

[0328] The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0329] Existing dance training systems struggle to be customized to individual user preferences and emotional states, failing to provide comprehensive analysis and intelligent feedback on users' real-time movements and psychological states, resulting in insufficient practice efficiency and emotional support. Furthermore, in group dance training, the lack of technology to automatically assign individual movements, collect multi-user data in real-time, and provide personalized and group-wide comprehensive feedback limits the improvement of overall group performance and collaboration.

[0330] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0331] In this invention, the server includes a device for receiving user behavior patterns, age information, physical abilities, preference information, and emotional information; a device for parsing user information and generating personalized motion videos and audios based on stored motion data and a generative artificial intelligence model; a device for sending the above content to a user terminal; and a device for parsing emotions and generating and providing adaptive guidance based on the analysis results. It also includes a device for real-time acquisition of motion data via electronic measurement devices, timely identification of problems through comparison with generated videos, generation of improvement suggestions for users based on these problems and emotional states, and guidance for users to correct their movements via the terminal. Furthermore, the server can receive and integrate information from various users in multi-user scenarios, automatically generate integrated group motion performance and assign individual motion tasks, acquire data from each member in real time, analyze overall and individual performance and emotions, and provide feedback to the group and individuals respectively. This enables intelligent generation of customized dance content based on users' personalized needs and emotional states, real-time feedback on both movement and emotion, and automation and efficiency of multi-user collaborative training, thereby significantly improving the effectiveness of dance training and personal emotional experience.

[0332] "Behavioral style" refers to information about the specific movement style, dance type, or exercise method that a user wishes to practice or perform.

[0333] "Age information" refers to relevant data that reflects the user's current actual age.

[0334] "Physical ability" refers to information describing a user's athletic ability, physical strength, flexibility, and other physical qualities.

[0335] "Preference information" refers to information about a user's personal preferences or inclinations regarding movement style, music genre, difficulty level, etc.

[0336] "Emotional information" refers to the information about the user's current psychological and emotional state as they subjectively feel during practice.

[0337] "Device" refers to hardware, software, or a combination thereof used to perform a specific function.

[0338] "Analysis" refers to the process of extracting, analyzing, and judging features from input information.

[0339] "Motion information set" refers to a data set that is pre-stored in a storage device and contains various motion patterns, posture parameters and motion characteristics.

[0340] "Generative artificial intelligence models" refer to artificial intelligence algorithms or systems that can automatically generate content (such as action videos or audio) based on input information.

[0341] "Personalized action videos" refer to exclusive action demonstration video content generated based on the user's individual characteristics and needs.

[0342] "Audio information" refers to audio content that matches the aforementioned action video and provides background or rhythm for the user's practice.

[0343] "Information terminal" refers to an electronic device that allows users to receive and display system outputs, such as a smartphone, tablet, or computer.

[0344] "Electronic measuring devices" refer to sensing devices that can be worn or equipped by users to collect and record motion data in real time, such as wristbands, ankle bracelets, and motion sensors.

[0345] "Technical issues" refer to deficiencies, errors, or areas for improvement discovered when comparing actual user actions with standard actions.

[0346] "Guidance and suggestions" refers to action improvement suggestions and emotional support prompts automatically generated by the system based on the user's actual performance and psychological state.

[0347] "Group integration action performance" refers to the team collaboration action content generated based on the individual information and cooperation relationships of multiple users.

[0348] "Individual sports division of labor" refers to the specific actions or performance tasks assigned to individual team members.

[0349] "Action performance" refers to the technical implementation and performance capability of a user or team when completing a specified action.

[0350] "Emotional state" refers to the psychological and emotional attitudes exhibited by users or teams as a whole or at specific moments during the training process.

[0351] This invention can be implemented in the following ways:

[0352] The server communicates and coordinates with information terminals (such as smartphones, tablets, and personal computers) and electronic measurement devices (such as wearable devices, fitness trackers, and motion sensors) via networks. The information terminals are equipped to collect user input and display dance content and feedback, which can be achieved through an app or web interface. The server-side is recommended to use high-performance computing hardware and be equipped with an operating system (such as Linux), a backend processing platform (such as Python + Django / Flask), a database (such as MySQL), and an artificial intelligence service framework (such as TensorFlow, PyTorch, OpenPose, and MediaPipe).

[0353] The server first receives information from the user, including: behavior patterns (such as dance type), age, physical abilities, preferences, and emotional state. The user fills in the corresponding information on the terminal interface, for example, by selecting options or typing. The terminal then sends this information to the server via secure protocols such as HTTPS.

[0354] The server processes data using a set of action information stored in a database and generative artificial intelligence models (such as those based on GPT, Stable Diffusion, and MusicLM). First, the server generates prompts, organizing the user input into structured text. For example:

[0355] Favorite dance style: Street dance

[0356] Current mood: Happy

[0357] Dance level: Intermediate

[0358] Age: 22

[0359] Preference: Music with a strong rhythm

[0360] The server takes the prompt as input and calls a generative artificial intelligence model to automatically generate personalized action videos (such as MP4 format) and matching audio information (such as MP3 format). The server stores the generated video and audio and directly sends the playback address or file to the information terminal.

[0361] Users download and play personalized motion videos and audio content on the terminal. They then practice the movements according to the video instructions. If the user is wearing a wearable device, the terminal collects various motion data of the user's body in real time via Bluetooth or other means (such as acceleration, joint angles, gait information, etc.), and can simultaneously record video using a camera to further improve the accuracy of motion analysis.

[0362] The terminal automatically uploads this motion data to the server. The server processes the actual motion data using motion recognition algorithms (OpenPose, MediaPipe, etc.) and compares it with standard personalized motion videos to automatically identify technical issues (such as uncoordinated movements, insufficient range of motion, etc.). The server also performs sentiment analysis on the emotional information input by the user, using existing NLP tools (such as BERT, FastText, etc.) to determine the user's possible psychological state.

[0363] The server analyzes the results and generates personalized guidance and suggestions (such as "Please increase the range of your arm swing" or "You're in a good mood, please keep it up"), which are then sent to the user's terminal. The terminal provides real-time prompts to the user via text, voice, and other means, helping them improve the accuracy of their movements and providing emotional support.

[0364] In multi-user collaborative scenarios, the server can receive information from multiple users, automatically generate group-integrated action designs based on generative artificial intelligence models, and assign specific individual movement tasks to each user. The server collects data from all members in real time, performs technical and emotional state assessments separately, and generates differentiated feedback for the team and individuals, achieving efficient group interaction and guidance.

[0365] For example, a user enters the following content into the terminal:

[0366] Favorite dance style: Jazz

[0367] Current emotion: tension

[0368] Dance level: Beginner

[0369] Age: 20

[0370] Preference: Soft music

[0371] The server generates jazz dance videos of appropriate difficulty with soft background music, and proactively pushes suggestions such as "Relax and practice slowly with the music" when it detects that the user is nervous.

[0372] Through the above technical solutions, the present invention can fully meet the diverse needs of users in personalized dance training, emotional support and group interaction, and ensure the feasibility and intelligence of the system.

[0373] use Figure 14 The processing procedure is explained.

[0374] Step 1:

[0375] Users open the application on their devices, enter personal information such as dance style, age, physical ability, preferences, and mood through the interface, and click the "Submit" button to send all the information to the server.

[0376] Input: The user's dance style, age, physical ability, preferences, and mood.

[0377] Output: Structured user information data.

[0378] Specific actions: The user selects a drop-down option, fills in the text box, and confirms submission.

[0379] Step 2:

[0380] The server receives user information data from the terminal and parses all input. Based on the parsing results, the server automatically generates corresponding prompt statements, structuring the user data into explicit prompt text for subsequent content generation.

[0381] Input: Structured user information uploaded by the terminal.

[0382] Output: Prompt statements suitable for generative artificial intelligence models.

[0383] Specific actions: The server extracts data fields and concatenates them into logically clear descriptive text.

[0384] Step 3:

[0385] The server inputs the generated prompts into a generative artificial intelligence model (such as a model for 3D motion video generation and AI music generation), and combines them with motion data from an internal database to automatically generate personalized dance video and audio content that matches the user's needs.

[0386] Input: Prompt statement, database action information.

[0387] Output: Personalized dance video files and audio files.

[0388] Specific actions: The server invokes the AI ​​model's inference function, combines database information with the generated results, and stores them.

[0389] Step 4:

[0390] The server pushes the storage address or file content of the generated video and audio files to the user's terminal via the network. The terminal receives and downloads the relevant multimedia files and plays them in a local player (such as a built-in video player).

[0391] Input: The address or content of the multimedia file returned by the server.

[0392] Output: Playable video and audio files on the terminal.

[0393] Specific actions: The terminal automatically initiates a download task, and after completion, it launches a media player to play the video.

[0394] Step 5:

[0395] Users watch and practice dance on the terminal while wearing wearable electronic measurement devices (such as smart bracelets). The terminal collects user motion data (such as acceleration, joint angles, etc.) in real time via Bluetooth and other means, and simultaneously captures video. The collected data is sent to the server in real time.

[0396] Input: Motion data collected by the user's wearing device, and video stream recorded by the camera.

[0397] Output: Real-time uploaded exercise data and practice videos.

[0398] Specific actions: The terminal automatically establishes a device connection, and the background continuously collects and uploads data.

[0399] Step 6:

[0400] After receiving motion data and video, the server calls motion recognition and analysis algorithms to compare the actual movements with standard dance moves generated by AI, and identifies deficiencies in the movements (such as insufficient range of motion or slow rhythm).

[0401] Input: motion data, recorded videos, personalized dance standard data.

[0402] Output: Problem analysis results after comparing the user's actual actions with the standard actions.

[0403] Specific actions: The server analyzes the data frame by frame and outputs the technical error points.

[0404] Step 7:

[0405] The server performs sentiment analysis on the emotional information obtained from user input or practice, and combines it with the results of action problem analysis to automatically generate personalized and highly adaptive action guidance suggestions and emotional support prompts.

[0406] Input: Emotional information and action analysis results.

[0407] Output: Personalized guidance and emotional feedback.

[0408] Specific actions: The server calls the sentiment analysis model, logically judges the suggested content, and organizes the feedback text.

[0409] Step 8:

[0410] The server pushes the generated technical improvement suggestions and emotional support feedback to the terminal. The terminal then reminds the user of the feedback in real time through interface pop-ups, voice broadcasts, and other means, helping them to adjust their practice methods and mindset in a timely manner.

[0411] Input: Feedback information sent by the server.

[0412] Output: The specific feedback content displayed or broadcast on the terminal.

[0413] Specific actions: After receiving a message, the terminal displays a pop-up window or calls the speech synthesis function to read the content.

[0414] Step 9:

[0415] Users can adjust their actions or mindset based on feedback from the terminal, and can repeatedly start the next round of practice and data collection, entering a new continuous iterative learning process to continuously improve their technical skills and emotional experience.

[0416] Input: Feedback information received and understood by the user.

[0417] Output: The user's new actions and states after repeated practice.

[0418] Specific actions: The user actively corrects their actions, watches the video repeatedly, and submits new collected data again through the terminal.

[0419] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0420] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0421] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0422] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0423] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0424] Second Implementation Method

[0425] Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0426] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0427] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0428] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0429] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0430] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to capture images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0431] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0432] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0433] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0434] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0435] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0436] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0437] Example 1

[0438] The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0439] Application Example 1

[0440] The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0441] Example 2

[0442] The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0443] Application Example 2

[0444] The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0445] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0446] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0447] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0448] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0449] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0450] Third Implementation Method

[0451] Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0452] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0453] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0454] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0455] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0456] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to capture images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0457] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0458] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0459] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0460] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0461] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0462] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0463] Example 1

[0464] The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0465] Application Example 1

[0466] The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0467] Example 2

[0468] The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0469] Application Example 2

[0470] The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0471] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0472] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0473] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0474] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0475] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0476] Fourth Implementation Method

[0477] Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0478] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0479] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0480] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0481] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0482] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0483] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0484] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0485] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0486] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0487] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0488] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0489] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0490] Example 1

[0491] The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0492] Application Example 1

[0493] The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0494] Example 2

[0495] The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0496] Application Example 2

[0497] The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0498] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input from the user representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0499] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be partially or entirely performed by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0500] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0501] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0502] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0503] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The system determines the user's emotions. Furthermore, the emotion-specific model 59 can similarly determine the robot's emotions, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0504] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0505] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0506] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0507] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0508] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0509] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0510] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0511] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0512] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0513] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0514] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0515] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0516] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0517] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0518] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0519] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0520] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0521] In addition, the following notes are provided in response to the above explanation.

[0522] Example 1

[0523] (Note 1)

[0524] An information processing system includes: an information input unit for inputting motion information, age information, physical ability information, and preference information from an information input device; a data conversion and transmission unit for converting user information obtained by the information input unit into a predetermined data format and transmitting it through a communication network; a generation unit for retrieving motion pattern information and audio information from an information storage device based on the user information, and inputting the information and prompt statements into a generative artificial intelligence model to generate personalized motion images and audio information; an information providing unit for sending the personalized motion images and audio information generated by the generation unit to a terminal device; and an output unit for displaying and playing the motion images and audio information by the terminal device.

[0525] (Note 2)

[0526] The information processing system according to Appendix 1 further includes: a data acquisition unit for acquiring user motion data in real time from a motion monitoring information acquisition device worn by the user; a data information analysis unit for comparing the user motion data acquired by the data acquisition unit with the personalized motion images generated by the generation unit, and analyzing motion differences and problems; an information generation unit for outputting personalized motion improvement guidance and motion adjustment feedback data based on the results of the information analysis unit; and an auxiliary unit for guiding the user to correct their motion in a physical or visual manner through the information acquisition device or terminal device based on the feedback data.

[0527] (Note 3)

[0528] The information processing system according to Appendix 1 further includes: an information integration and allocation unit for receiving personalized action information or preference information from multiple users, generating integrated action performance information based on group identification information, and simultaneously generating and allocating individual task action information and audio information for each user; an analysis and evaluation unit for acquiring action data of each user in real time, analyzing the action content of the group as a whole and each user, and evaluating group coordination and individual action accuracy; and an information provision unit for generating and distributing performance improvement suggestions, guidance, or feedback data information by group and individual based on the results of the analysis and evaluation unit.

[0529] Application Example 1

[0530] (Note 1)

[0531] An information processing system includes: a device for acquiring environmental information including user body movement information and preference conditions via a terminal; a device for parsing the environmental information and body movement information sent from the terminal, and inputting predetermined prompt information into a generative artificial intelligence model to generate an optimal movement pattern, body movement content, and sound information; a device for sending the generated movement pattern, body movement content, and sound information to an output device or terminal; a device for acquiring actual movement data of the user or mechanical device in real time and sending it to a server; a device for analyzing the real-time acquired movement data and the generated information to generate evaluation and feedback information; and a device for sending the feedback information to the terminal or mechanical device to guide movement correction or improvement.

[0532] (Note 2)

[0533] According to the information processing system described in Appendix 1, a wearable body information acquisition device is used to acquire motion signals and physiological signals from multiple body parts in real time, and provide them to the generative artificial intelligence model and analysis device to simultaneously analyze individual state and emotional state, and provide adaptive guidance and improvement information.

[0534] (Note 3)

[0535] The information processing system described in Appendix 1 is used to acquire body movement information, preference conditions, and emotional states from multiple users through a terminal, generate collaborative movement patterns, individual allocation instructions, and audio information based on a generative artificial intelligence model, and output personalized optimal information to each user terminal or mechanical device. At the same time, it analyzes group movement data and individual movement data to provide evaluation, guidance, and feedback information for individuals and groups.

[0536] Example 2

[0537] (Note 1)

[0538] An information processing system includes: a device for acquiring information on action type, age group, physical ability, preferences, and emotional state from a user; a device for parsing the acquired information and, based on information storage on the information processing device, generating action video data and audio data based on user information by inputting prompt data through a generative artificial intelligence model; a device for transmitting the generated action video data and audio data to an information display device; and a device for parsing the output content of the generative artificial intelligence model and adding control signals according to user preferences and emotional state.

[0539] (Note 2)

[0540] The information processing system according to Appendix 1 further includes: means for acquiring temporal motion information from a wearable input device; means for parsing the acquired temporal motion information and comparing it with generated motion video data to automatically identify discrepancy areas or problem parts; means for identifying the user's emotional state through an emotion analysis device and generating adaptive guidance content and response information based on the problem parts and emotional state; and means for outputting control signals to the wearable input device and providing physical guidance for motion correction through vibration or other means.

[0541] (Note 3)

[0542] The information processing system according to Appendix 1 further includes: a device for acquiring action type, preference, and emotional state information from multiple users, and performing action integration processing based on such information to generate collective action performance data; a device for distributing and transmitting the assigned action composition information to multiple users; a device for acquiring action timing data from multiple users and performing action evaluation processing on the group and individuals; and a device for generating and outputting guidance content and response information for the group and individuals based on the collective and individual action evaluation processing results and emotional analysis results.

[0543] Application Example 2

[0544] (Note 1)

[0545] An information processing system includes: means for receiving behavioral patterns, age information, physical abilities, preference information, and emotional information from a user; means for parsing the received information and generating personalized action video and audio information based on a set of action information stored in a storage device on a computing device and a generative artificial intelligence model; means for sending the generated personalized action video and audio information to an information terminal; and means for parsing the emotional information and generating adaptive guidance and suggestion information based on the parsing results and providing it to the information terminal.

[0546] (Note 2)

[0547] The information processing system according to Appendix 1 further includes: a device for acquiring motion information in real time through an electronic measuring device worn by the user; a device for parsing the acquired motion information and comparing it with the generated personalized motion video to extract technical problems; a device for generating guidance and suggestion information to support user motion improvement and mood improvement based on the extracted technical problems and emotion analysis results, and providing it to an information terminal; and a device for presenting auxiliary information to the user through the electronic measuring device to guide their motion correction.

[0548] (Note 3)

[0549] The information processing system according to Appendix 1 further includes: a device for receiving individual behavioral patterns, preference information, and emotional information from multiple users and generating a group-wide integrated action performance; a device for assigning corresponding individual motor tasks to each user; a device for acquiring the motor and emotional information of each user in real time and analyzing the action performance and emotional state of the group as a whole and individuals; and a device for providing guidance and suggestions to the group as a whole and each user to support action improvement and emotional optimization.

Claims

1. An information processing system, characterized in that, include: Device for receiving information including a user’s dance style, age, physical strength and preferences; A device used to analyze the received information and generate original dance videos and music based on existing dance databases; and A device used to send the generated original dance videos and music to a user's terminal.

2. The information processing system according to claim 1, characterized in that, Also includes: A device for acquiring motion data in real time through wearable devices worn by users; A device for analyzing the acquired motion data and comparing it with the generated original dance video to identify problem areas; Devices for providing guidance and feedback to users based on identified problem points; and A device for guiding user movement corrections via the wearable device.

3. The information processing system according to claim 1, characterized in that, Also includes: A device used to receive personalized dance style and preference information from multiple users and generate a comprehensive dance performance as a group. A device used to assign individual dance roles to each member; A device for acquiring and analyzing the movement data of each member in real time; and A device used to provide feedback to the group as a whole and to individual members, and to help improve performance.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A