system

A system that collects, standardizes, and personalizes exercise videos with user feedback enhances motivation and cognitive function in the elderly by combining failure and success scenes, optimizing visual and auditory elements.

JP2026071547APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Conventional methods fail to effectively motivate the elderly to exercise, leading to a decline in cognitive function and lack of sustained interest.

Method used

A system that collects and standardizes exercise video material, generates personalized videos combining failure and success scenes, optimizes visual and auditory elements, and continuously improves based on user feedback to enhance motivation.

Benefits of technology

The system effectively stimulates the elderly's willingness to exercise, forming spontaneous habits and maintaining cognitive function by providing tailored and engaging content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071547000001_ABST
    Figure 2026071547000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 Means for collecting motion video materials, Means for analyzing and standardizing the collected motion video materials, Means for generating a video scenario based on the standardized video materials, Means for generating an image using the generated video scenario, Means for editing the generated image and optimizing visual and auditory elements, Means for distributing the edited and optimized image to users, Means for collecting and analyzing feedback from users, Means for improving the video generation model based on the analysis results, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Lack of motivation for exercise among the elderly and the accompanying decline in cognitive function are recognized as social problems. Conventional approaches have been difficult to provide effective motivation for this problem. There is a need for a new method to stimulate the elderly's willingness to exercise voluntarily.

Means for Solving the Problems

[0005] This invention provides a system for increasing exercise motivation among the elderly. This system collects and standardizes exercise video material, then generates videos that evoke a feeling in viewers of wanting to perform the exercise themselves. The generated videos are edited and optimized, and delivered to users in an individualized format. Furthermore, it includes means for continuously improving the video generation model by collecting and analyzing user feedback. This makes it possible to encourage the elderly to develop an interest in exercise, form spontaneous exercise habits, and ultimately support the maintenance of cognitive function.

[0006] "Exercise video footage" refers to video data demonstrating exercises and workouts suitable for the elderly, including scenes of both failure and success.

[0007] "Analysis" is the process of examining the format, resolution, and length of collected material to ensure consistency.

[0008] "Standardization" is the process of unifying data formats so that materials can be used in generative models.

[0009] A "video script" is the planning of the arrangement and sequence of materials to evoke a specific emotion or action in the viewer.

[0010] "Generating video" is the process of creating new video content that evokes emotions in viewers, based on a planned scenario.

[0011] "Optimization" refers to editing and adjusting the generated video so that it appeals to the viewer both visually and aurally.

[0012] "Distributing to users" refers to sending the generated video to specific users and making it available for viewing.

[0013] "Collecting and analyzing feedback" means using a system to gather user reactions and post-viewing comments, evaluate them, and use that information to make improvements.

[0014] "Improving the video generation model" is the process of adjusting or evolving the video generation algorithm based on feedback. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention is a system for increasing exercise motivation, primarily targeting the elderly. This system effectively functions through a series of processes utilizing a server and terminals. The server first collects exercise video material for the elderly from a database, analyzes it, and standardizes it. The standardized material is selected based on scenarios designed to evoke specific emotions.

[0037] The server operates a video generation model based on these selected scenarios to generate new footage. This footage combines scenes of deliberate failure with scenes showing the possibility of success, designed to motivate viewers to take on the challenge themselves. The generated footage is edited and visually and aurally optimized. It is then delivered to the device in a personalized format, adjusted based on the user's interests and past viewing history.

[0038] The feelings and experiences that users have after watching the videos and actually engaging in the exercises are collected as feedback. The server analyzes this feedback and stores it as data to help generate the next videos. Through this analysis process, the system is continuously improved, enhancing its ability to more effectively motivate viewers to exercise.

[0039] For example, a user might watch a video streamed on their device and be inspired to try stretching themselves. This user's data is collected, and their post-viewing feedback and the extent of their exercise are analyzed. Based on this analysis, more refined content tailored to that user can be delivered in the future. This method encourages users to continue exercising enthusiastically, which in turn contributes to maintaining and improving cognitive function.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] The server collects exercise videos for seniors from a database. This material includes both fully successful and intentionally unsuccessful exercises.

[0043] Step 2:

[0044] The server analyzes the collected video footage, unifying and standardizing its resolution and format. This prepares the video generation model for smooth processing.

[0045] Step 3:

[0046] The server designs scenarios that combine failure and success scenes from video footage to effectively evoke emotions in viewers. These scenarios are built based on past viewing data.

[0047] Step 4:

[0048] The server executes a video generation model based on the selected scenario and generates new footage. This footage is structured to make viewers think, "I want to do this myself."

[0049] Step 5:

[0050] The server edits the generated video and optimizes it by adding visual and auditory elements. This includes background music and text messages to enhance the video's appeal.

[0051] Step 6:

[0052] The device delivers the completed video to the user. The video is delivered in a format that is individually tailored to the user's interests and viewing history.

[0053] Step 7:

[0054] Users watch the streamed videos and, inspired by them, try exercising themselves. Their impressions after watching the videos and the results of their exercise are collected as feedback.

[0055] Step 8:

[0056] The server analyzes user feedback and stores data to help generate future videos. This allows the system to continuously improve and provide more effective video content.

[0057] (Example 1)

[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0059] Maintaining sustained exercise motivation among older adults is crucial for preserving their physical and cognitive health. However, traditional exercise programs are generally uniform and fail to adequately consider individual motivations and interests. Furthermore, there is a lack of mechanisms for effectively utilizing user feedback, making continuous improvement difficult. This leads to the challenge that older adults tend to lose interest in exercise.

[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0061] In this invention, the server includes means for collecting exercise content, means for analyzing and standardizing the collected exercise content, and means for generating scenarios based on the standardized content. This makes it possible to stimulate the user's interest and emotions through personalized video and continuously increase their motivation to exercise.

[0062] "Exercise content" refers to media materials such as videos and audio used to promote physical activity for the elderly.

[0063] "Standardization" is the process of converting motion content from different formats and with different content into a consistent format, making it easier to analyze.

[0064] A "scenario" is the flow and structure of action content designed to evoke specific emotions or motivations in the viewer.

[0065] "Media" is a general term for files and streams containing visual and auditory elements that are generated for distribution to users.

[0066] A "generative model" is an algorithm or system that automatically generates new content based on specific prompts.

[0067] "Editing" is the process of modifying the visual and auditory elements of generated media in order to provide an optimal user experience.

[0068] "Personalization" refers to optimizing the media delivered to each user based on their past usage history and interests.

[0069] "Feedback" refers to the impressions, opinions, and experience-based information that users provide after viewing, and it is data that helps improve the system.

[0070] This invention is a system for improving the exercise motivation of the elderly, and it functions through the cooperation of the server, terminal, and user.

[0071] The server collects exercise content from a database, analyzes it, and standardizes it. This is done using video recognition software and analysis tools. For example, it is possible to analyze video material and extract features using Python's OpenCV or TENSORFLOW®. Based on the analyzed data, the server generates exercise scenarios tailored to the elderly and creates new media based on a generative AI model. The generative AI model used here is a general algorithm that processes text and image prompts.

[0072] The generated media is edited and optimized by the server. This editing process uses video editing software (e.g., Adobe Premiere Pro) to make visual and auditory adjustments to improve the quality of the video. Specifically, this includes adjusting brightness and compositing audio.

[0073] Optimized media is sent to the device and personalized based on the user's past usage history and interests. The device uses machine learning models to analyze the user's profile and provide optimal content. This suggests content tailored to each user, maintaining a high level of motivation for exercise.

[0074] Users watch the provided video and perform exercises according to the instructions. They can then input feedback based on their experience into their device. The server collects this feedback and uses it to continuously improve the system's accuracy and quality for future media generation.

[0075] As a concrete example, by inputting a prompt message into the AI ​​model such as, "Please generate videos that seniors can enjoy and that will increase their motivation to exercise," it is possible to generate corresponding media. Through this process, viewers can continuously increase their motivation to exercise.

[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0077] Step 1:

[0078] The server collects exercise content for seniors from a database. The input is metadata and URLs of exercise videos, and the output is raw data for analysis and standardization. Specifically, it queries the database via an API and retrieves the corresponding video files.

[0079] Step 2:

[0080] The server analyzes and standardizes the collected motion content. Raw data obtained in step 1 is passed as input, and standardized data is generated as output. For data processing, important movements within the video are extracted using Python's OpenCV, and these movements are analyzed using TensorFlow. Specifically, each video frame is analyzed, and motion features are extracted as numerical data.

[0081] Step 3:

[0082] The server generates scenarios based on standardized data. The input is the output data from step 2. The output is a scenario designed to evoke emotions in the viewer. Specifically, it uses an emotion analysis algorithm to analyze which scenes evoke which emotions and constructs the optimal scenario.

[0083] Step 4:

[0084] The server generates media using a generative AI model based on the generated scenario. The input consists of a scenario and a prompt (e.g., "Please generate a video that elderly people can enjoy and that will motivate them to exercise"), and the output is the creation of new media. Specifically, the prompt is input to the generative AI model, and the AI ​​automatically outputs media combining video clips and audio.

[0085] Step 5:

[0086] The server edits the generated media and optimizes its visual and auditory elements. The input is the output media from step 4. The output is the visually and aurally optimized media. Specifically, Adobe Premiere Pro is used to edit video transitions and background music to provide a pleasant viewing experience.

[0087] Step 6:

[0088] The device receives optimized media and personalizes it based on the user's past usage history and interests. Viewing history data and optimized media are provided as input. Personalized media is generated as output. Specifically, the device uses a machine learning model to analyze the user's profile and select and deliver the most relevant content.

[0089] Step 7:

[0090] Users watch the delivered media, perform the exercise, and then input feedback into the device. The input consists of the user's experience and impressions, and the output is feedback data. Specifically, the user enters their impressions and suggestions for improvement into a form provided on the device and submits it to the system.

[0091] Step 8:

[0092] The server analyzes the collected feedback and uses it to improve future media generation. User feedback data is provided as input. An improved generative model is obtained as output. Specifically, a natural language processing library is used to analyze the feedback text, classify positive and negative emotions, and reflect these in the next content generation.

[0093] (Application Example 1)

[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0095] There is a challenge in the lack of methods to sustainably increase the motivation for exercise among the elderly. Conventional methods of promoting exercise are insufficient in providing personalized content to maintain interest and in generating adaptive content based on feedback.

[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0097] In this invention, the server includes means for collecting exercise media material, means for analyzing and standardizing the collected exercise media material, and means for generating media scenarios based on the standardized data material. This makes it possible to provide appropriately personalized exercise recommendation content to individual elderly people.

[0098] "Movement media material" refers to all digital content that uses movement as its theme and is provided through visual and auditory means.

[0099] "Means of analysis and standardization" refers to techniques for analyzing collected material data and organizing it into a consistent quality and format.

[0100] A "media scenario" refers to a storyline and structure designed to evoke emotions and interest in the audience.

[0101] "Means of generating information" refers to technologies for creating new visual and auditory content based on media scenarios.

[0102] "Means of optimizing visual and auditory elements" refers to technologies that adjust the sound and visuals of content to improve the user experience.

[0103] "Means of distribution to users" refers to the methods and technologies used to deliver generated viewing content to individual users in an appropriate format.

[0104] "Means of collecting and analyzing feedback" refers to technologies that collect user reactions and data and analyze them to help in creating future content.

[0105] "Means of improving information generation models" refers to technologies for improving the content generation process based on collected feedback.

[0106] "Personalized exercise recommendation content" refers to digital content about exercise that is customized based on each user's interests and activity history.

[0107] The system for implementing this invention mainly consists of a server and terminals. The server collects and analyzes exercise media material and generates media scenarios based on standardized data. Specifically, the server optimizes visual and auditory content using platforms such as Google® Cloud AI and Amazon SageMaker. This creates new information that is of interest to the user and delivers it to the terminals. The terminals are smartphones, tablets, etc., and play the role of delivering personalized content to the user. Users can view this content and provide feedback. The server collects and analyzes user feedback to improve future content generation models. This provides more personalized exercise recommendation content and helps to stimulate the user's motivation to continue exercising.

[0108] For example, if a user watches content titled "Morning stretches in the park," a new scenario, such as "Stretches in a spring park with cherry blossoms," will be generated based on their reaction. An example of a prompt for the generating AI is, "Generate a morning stretching video for seniors, using a relaxed park scene."

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] The server collects sports media material from a database. This material includes video and audio files containing specific sports scenes. For input, it sends queries to the database based on pre-configured acquisition conditions and filter information, and receives all the material. The output is a list of the material stored on the server. Specifically, it uses a Python library to retrieve material from the database and converts it into a format that can be managed on the server.

[0112] Step 2:

[0113] The server analyzes and standardizes the collected motion data. During the analysis phase, it extracts information such as the type, duration, and difficulty of the motion contained in the data. The input is the data accumulated in step 1, and the output is structured metadata. Specifically, it uses a machine learning model to classify the data and extract specific features.

[0114] Step 3:

[0115] The server generates media scenarios based on standardized data. These scenarios include a narrative structure designed to engage the audience. The inputs are the standardized data obtained in step 2 and specific scenario generation rules, and the output is a completed media scenario. Specifically, it generates a storyline constructed using natural language processing.

[0116] Step 4:

[0117] The server generates new visual and auditory content based on the generated scenario. It utilizes a generative AI model to create visual materials for the target scene. The input is the scenario and template information generated in step 3, and the output is high-quality video and audio files. Specifically, prompt text is input to the AI ​​model to generate video content corresponding to the scenario.

[0118] Step 5:

[0119] The server optimizes the generated content visually and aurally. This optimization includes color correction and volume adjustment. The input is the content obtained in step 4, and the output is the visually and aurally optimized content. Specifically, video editing software is used to improve video quality through automatic correction functions.

[0120] Step 6:

[0121] The server delivers optimized content to the terminal. This delivery is done in real time and is adjusted to the individual user's preferences. The input is the content prepared in step 5, and the output is the video stream sent to the user's terminal. Specifically, data transfer is performed using streaming technology.

[0122] Step 7:

[0123] Users view the provided content and provide feedback. This feedback records their impressions after viewing the content and details of the exercises they actually performed. The input is the user's reaction to viewing the content, and the output is structured feedback data. Specifically, data is entered through the user interface using web forms and questionnaires.

[0124] Step 8:

[0125] The server analyzes the collected feedback data and uses it to improve the content it generates next. The input is the feedback data collected in step 7, and the output is the content generation model updated based on the feedback. Specifically, it uses data analysis tools to statistically analyze the feedback and uses it as training data for the AI ​​model.

[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0127] This invention combines an emotional engine with a system designed to enhance exercise motivation, and is primarily aimed at the elderly. The system consists of a server and terminals and functions through a process that effectively stimulates the user's desire to exercise.

[0128] The server first collects video material from a database, analyzes it, and standardizes it. This video material includes exercises and challenges that are easy for seniors to participate in. Next, based on the collected data, the server generates scenarios that make viewers think, "I want to try that myself," and creates new videos using a video generation model.

[0129] The generated video is edited, and its visual and auditory elements are optimized. At this stage, the device delivers the video to the user and recognizes the user's emotions in real time through an emotion engine. The emotion engine analyzes the user's facial expressions and body movements to determine what emotions they are experiencing. Based on this information, the device instantly adjusts the video content to keep the user as engaged as possible.

[0130] Users move their bodies while watching the streamed videos. During this time, they can input their feelings and post-exercise feedback through their device. This feedback is sent to the server and stored in a database along with the analysis results obtained by the emotion engine. The server uses this information to improve future video generation and emotion recognition, thereby adjusting the system to enhance its overall effectiveness.

[0131] For example, if a user smiles while watching a video, the emotion engine recognizes this, and the device switches to displaying a video with more enjoyable elements. In this way, users enjoy the process of exercising, and an exercise habit is naturally formed. Furthermore, the level of exertion the user is experiencing is recorded through feedback, allowing for the optimization of content to suit individual needs.

[0132] This system, equipped with a dynamic content adjustment engine based on emotions, can strengthen the motivation of older adults to engage in exercise spontaneously and contribute to maintaining cognitive function.

[0133] The following describes the processing flow.

[0134] Step 1:

[0135] The server extracts exercise videos for seniors from its database. These videos include a wide variety of content, from simple exercises to scenes of deliberate failure.

[0136] Step 2:

[0137] The server analyzes the collected material, standardizing it by unifying the format and adjusting the resolution as needed. This standardization makes it suitable as input for the generative model.

[0138] Step 3:

[0139] The server generates scenarios based on the user's past viewing history and preferences. These scenarios effectively combine scenes of failure and scenes of potential success.

[0140] Step 4:

[0141] The server uses the generated scenario to create new video content using a video generation model. This generation process is designed to motivate the viewer.

[0142] Step 5:

[0143] The server edits the generated video content and optimizes its visual and auditory elements by adding music and text messages, thereby enhancing the emotional appeal of the video.

[0144] Step 6:

[0145] The device delivers edited video to the user. Equipped with an emotion engine, it analyzes the user's facial expressions and movements in real time and recognizes their emotions.

[0146] Step 7:

[0147] The emotion engine analyzes the user's emotions while viewing content and provides the results to the device, instantly adjusting the video content as needed. This ensures that content tailored to the user is continuously displayed.

[0148] Step 8:

[0149] Users provide their emotions and reactions to the movement to the device, which are recorded as feedback after viewing the video. This feedback is then collected.

[0150] Step 9:

[0151] The server analyzes the sentiment analysis results and user feedback sent from the terminal and uses this information to improve the next video generation process. This enhances the overall efficiency of the system.

[0152] (Example 2)

[0153] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0154] In modern society, the impact of lack of exercise on the health of older adults is serious. Traditional methods of promoting exercise are not always suitable for all older adults, and problems such as a lack of motivation and inability to adapt to individual needs have emerged. In particular, there are situations where it is difficult to effectively stimulate the desire to exercise.

[0155] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0156] In this invention, the server includes means for collecting exercise-related data, means for analyzing and standardizing the collected exercise-related data, and means for generating video scenarios based on the standardized data. This enables dynamic exercise promotion that responds to the user's interests and emotions.

[0157] "Exercise-related data" refers to information related to the exercise a user performs, including the type, intensity, and frequency of exercise.

[0158] "Analysis" refers to the process of thoroughly examining collected data and transforming its contents into an understandable format.

[0159] "Standardization" refers to the process of organizing data from different formats and standards according to a unified set of criteria.

[0160] A "video scenario" refers to a plan or storyline for video production, which comprises content that will interest the user.

[0161] "Editing" refers to the process of modifying video and audio to create a final work.

[0162] "Visual and auditory elements" refer to all the components in a video that a user can perceive visually and aurally.

[0163] "Emotional recognition" refers to technology that uses sensors to read a user's facial expressions and actions and uses that information to determine their emotional state.

[0164] "Dynamic adjustment" refers to the process of receiving information in real time and changing content or its elements in response to that information.

[0165] "Response" refers to the reaction or feedback that a user inputs into the system.

[0166] "Analyzed response" refers to information obtained as a result of the system analyzing user feedback.

[0167] "Personalization" refers to the process of tailoring the content provided to each user to their individual interests and history.

[0168] "Observation history" refers to records of what a user has watched in the past and their viewing trends.

[0169] This invention is a system designed to increase the motivation of elderly and other users to exercise. This system works in conjunction with a server and a terminal to provide users with personalized exercise videos.

[0170] The server collects exercise-related data from a database. This includes video footage of simple exercises suitable for seniors, as well as some with a slight challenge element. The server analyzes this data and standardizes and unifies the format. Video editing software is used in this process to organize the data format and make it reusable.

[0171] Next, the server uses a generative AI model to generate video scenarios that will capture the viewer's interest. This involves using prompts, such as "Create a scene of light exercise for seniors to enjoy," to input instructions into the generative AI model. Based on this scenario, the newly generated video is adjusted, and its visual and auditory elements are optimized. Specifically, CG technology is used to prepare exercise scenes suitable for the user.

[0172] The device delivers the generated video to the user. When the user watches the delivered video, the emotion engine recognizes the user's emotions in real time. This recognition uses technology that reads the user's facial expressions and body movements using cameras and sensors. For example, if the user smiles, the emotion engine will determine that the user is "enjoying" the expression, and the device will automatically adjust the video content based on that feedback.

[0173] Users move their bodies while watching these videos, and after exercising, they input feedback into their device. This feedback is sent to a server and stored in a database for analysis. The analyzed feedback information is used to improve future video generation and optimize the entire system.

[0174] This system allows users to exercise while enjoying videos tailored to their interests and physical condition, which is expected to increase their motivation to exercise and help them establish regular exercise habits.

[0175] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0176] Step 1:

[0177] The server collects exercise-related data from a database. This input data includes video materials of exercises that are easy for seniors to participate in. The server classifies these materials and selects content suitable for the target users. Specifically, videos containing simple exercises such as stretching and walking are extracted.

[0178] Step 2:

[0179] The server analyzes and standardizes the collected motion-related data. The analysis process corrects inconsistencies in video format and resolution, generating a unified format. The input to this process is the collected raw data, and the output is reusable, standardized data. Editing software is used to organize the data in different formats.

[0180] Step 3:

[0181] The server uses a generative AI model to generate video scenarios based on standardized data. During this process, the prompt "Create a scene of light exercises for seniors to enjoy" is input, and the AI ​​generates the scenario. The output is scenario data containing content that seniors would find interesting, and this is used in the next step to generate the video.

[0182] Step 4:

[0183] The server generates new video using the generated scenario data. At this stage, CG technology is used to create exercise scenes for the elderly. The input is scenario data, and the output is generated video for distribution to users. The generated video has enhanced visual and auditory elements.

[0184] Step 5:

[0185] The terminal delivers the generated video to the user. Streaming technology is used for the user's device to ensure smooth playback. The input is the generated video from the server, and the output is the user's viewing screen.

[0186] Step 6:

[0187] The emotion engine built into the device recognizes the user's emotions in real time. Cameras and sensors capture the user's facial expressions and body movements, providing them as input information. The output of the emotion analysis is data indicating the user's emotional state, which is used for video adjustment.

[0188] Step 7:

[0189] The device dynamically adjusts the video content based on the output of the emotion engine. For example, if the user smiles to indicate enjoyment, more cheerful scenes will be added to the video. The input is emotion data, and the output is the adjusted video displayed to the user.

[0190] Step 8:

[0191] Users enter feedback into a device after their workout. This feedback includes their workout experience and impressions. This feedback is sent from the device to a server and used for analysis.

[0192] Step 9:

[0193] The server integrates feedback and analysis results from the emotion engine to optimize the system. The input to this process is user feedback data, and the output is improved system specifications for future video generation and content adjustments.

[0194] (Application Example 2)

[0195] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0196] In industrial machinery, reducing the mental and physical burden on users and maintaining sustained motivation are crucial for supporting efficient work execution and improving productivity. However, conventional methods have made it difficult to provide appropriate feedback and adjustments based on the machine's condition and the user's mental state. This leads to decreased work efficiency and, consequently, economic losses, posing a significant challenge.

[0197] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0198] In this invention, the server includes means for collecting motion data material, means for analyzing and standardizing the collected motion data material, and means for generating information scenarios based on the standardized data material. This makes it possible to recognize and analyze the user's emotions in real time and dynamically adjust the machine's work content and pace accordingly.

[0199] "Motion data material" refers to data related to machine operation information and user work movements, and is used to propose efficient work processes.

[0200] "Analysis and standardization" is the process of converting collected data into a basic format and preparing the data necessary to generate new information.

[0201] An "information scenario" is a set of instructions for a series of actions or reactions, generated based on analyzed and standardized data.

[0202] "Expression" refers to content that is generated using an information scenario and presented to the user through visual and auditory elements.

[0203] "Emotional state" refers to the current psychological and emotional condition of the user, based on their facial expressions and actions.

[0204] "Dynamic adjustment" means changing the content and pace in real time according to the user's emotional state.

[0205] "Feedback" refers to responses and information provided based on the user's experiences and evaluations, and is used as material for improving the system.

[0206] To realize this invention, a server-centric system is built, starting with collecting motion data, analyzing it, and standardizing it. Based on the collected data, the server generates information scenarios and uses these scenarios to form representations. The generated representations are edited to optimize visual and auditory elements and delivered to users via terminals.

[0207] The server utilizes cameras and sensors mounted on the machine to recognize the user's emotional state in real time and analyze the user's facial expressions and movements. Specific examples include the use of emotion analysis libraries such as OpenCV and dlib. This allows for dynamic adjustment of the content of the expressions in response to the user's psychological reactions.

[0208] Users can adjust their work pace and content through the delivered expressions, thereby improving efficiency. For example, if fatigue is detected while a user is performing a certain task, the server can provide expressions intended to promote relaxation, thereby improving the user's work efficiency.

[0209] Furthermore, user feedback is sent to the server via the device. This feedback information will be used to improve expression generation and emotion recognition technologies in the future. Through this iterative learning process, the accuracy of the generative AI model will be improved.

[0210] An example of a prompt message might be, "If the robot shows signs of fatigue during work, generate a refresh video and provide a positive scenario to encourage a break." This allows the server to generate an appropriate scenario based on the instructions, creating a dynamic work environment.

[0211] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0212] Step 1:

[0213] The server collects motion data material through various sensors and cameras. This input data includes videos of work movements and environmental conditions such as temperature. The server converts this data into a digital format and stores it in a database.

[0214] Step 2:

[0215] The server analyzes and standardizes the collected motion data. It uses raw data from the database as input and performs data processing such as noise reduction and data transformation. The standardized data forms the basis for the next processing step.

[0216] Step 3:

[0217] The server generates information scenarios based on standardized data. Utilizing a generative AI model, the algorithm develops the optimal scenario based on the prompt. This output is saved as an operational program and used in the next step.

[0218] Step 4:

[0219] The server generates visual and auditory representations using the generated information scenarios. It performs data calculations to create an effective combination of video and audio, and outputs the results as representation data.

[0220] Step 5:

[0221] The server edits and optimizes the generated representation data. Data processing, such as filtering and adding effects, is performed to balance the visual and auditory elements. This optimized output is then ready to be presented to the user.

[0222] Step 6:

[0223] The terminal delivers optimized content to the user. The user views the presented content and uses it as a guide for their work. Because this delivery is done in real time, a network protocol for delivery is used.

[0224] Step 7:

[0225] The device recognizes the user's emotional state in real time using a camera and sensors. The data obtained as input is analyzed using libraries such as OpenCV to identify the emotional state. The analysis results are then sent to a server.

[0226] Step 8:

[0227] The server dynamically adjusts the content of its expressions based on sentiment analysis results. Using prompt sentences as a guide, the generative AI model updates the scenario and outputs appropriate content in real time to keep the user interested.

[0228] Step 9:

[0229] Users input subjective experiences and evaluations into their devices through feedback. This input data is sent to a server for improvement and used in future representation generation models.

[0230] Step 10:

[0231] The server uses the collected feedback and emotion recognition results to improve its expression generation model. By analyzing input data and training the generative AI model, it enhances the accuracy of future scenario generation.

[0232] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0233] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0234] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0235] [Second Embodiment]

[0236] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0237] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0238] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0239] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0240] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0241] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0242] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0243] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0244] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0245] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0246] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0247] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0248] This invention is a system for increasing exercise motivation, primarily targeting the elderly. This system effectively functions through a series of processes utilizing a server and terminals. The server first collects exercise video material for the elderly from a database, analyzes it, and standardizes it. The standardized material is selected based on scenarios designed to evoke specific emotions.

[0249] The server operates a video generation model based on these selected scenarios to generate new footage. This footage combines scenes of deliberate failure with scenes showing the possibility of success, designed to motivate viewers to take on the challenge themselves. The generated footage is edited and visually and aurally optimized. It is then delivered to the device in a personalized format, adjusted based on the user's interests and past viewing history.

[0250] The feelings and experiences that users have after watching the videos and actually engaging in the exercises are collected as feedback. The server analyzes this feedback and stores it as data to help generate the next videos. Through this analysis process, the system is continuously improved, enhancing its ability to more effectively motivate viewers to exercise.

[0251] For example, a user might watch a video streamed on their device and be inspired to try stretching themselves. This user's data is collected, and their post-viewing feedback and the extent of their exercise are analyzed. Based on this analysis, more refined content tailored to that user can be delivered in the future. This method encourages users to continue exercising enthusiastically, which in turn contributes to maintaining and improving cognitive function.

[0252] The following describes the processing flow.

[0253] Step 1:

[0254] The server collects exercise videos for seniors from a database. This material includes both fully successful and intentionally unsuccessful exercises.

[0255] Step 2:

[0256] The server analyzes the collected video footage, unifying and standardizing its resolution and format. This prepares the video generation model for smooth processing.

[0257] Step 3:

[0258] The server designs scenarios that combine failure and success scenes from video footage to effectively evoke emotions in viewers. These scenarios are built based on past viewing data.

[0259] Step 4:

[0260] The server executes a video generation model based on the selected scenario and generates new footage. This footage is structured to make viewers think, "I want to do this myself."

[0261] Step 5:

[0262] The server edits the generated video and optimizes it by adding visual and auditory elements. This includes background music and text messages to enhance the video's appeal.

[0263] Step 6:

[0264] The device delivers the completed video to the user. The video is delivered in a format that is individually tailored to the user's interests and viewing history.

[0265] Step 7:

[0266] Users watch the streamed videos and, inspired by them, try exercising themselves. Their impressions after watching the videos and the results of their exercise are collected as feedback.

[0267] Step 8:

[0268] The server analyzes user feedback and stores data to help generate future videos. This allows the system to continuously improve and provide more effective video content.

[0269] (Example 1)

[0270] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0271] Maintaining sustained exercise motivation among older adults is crucial for preserving their physical and cognitive health. However, traditional exercise programs are generally uniform and fail to adequately consider individual motivations and interests. Furthermore, there is a lack of mechanisms for effectively utilizing user feedback, making continuous improvement difficult. This leads to the challenge that older adults tend to lose interest in exercise.

[0272] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0273] In this invention, the server includes means for collecting exercise content, means for analyzing and standardizing the collected exercise content, and means for generating scenarios based on the standardized content. This makes it possible to stimulate the user's interest and emotions through personalized video and continuously increase their motivation to exercise.

[0274] "Exercise content" refers to media materials such as videos and audio used to promote physical activity for the elderly.

[0275] "Standardization" is the process of converting motion content from different formats and with different content into a consistent format, making it easier to analyze.

[0276] A "scenario" is the flow and structure of action content designed to evoke specific emotions or motivations in the viewer.

[0277] "Media" is a general term for files and streams containing visual and auditory elements that are generated for distribution to users.

[0278] A "generative model" is an algorithm or system for automatically generating new content based on a specific prompt.

[0279] "Editing" is a process that modifies the visual and auditory elements of the generated media to provide an optimal user experience.

[0280] "Personalization" means optimizing the media to be delivered for each user based on the user's past usage history and interests.

[0281] "Feedback" is information based on the impressions, opinions, and experiences provided by the user after viewing, and is data useful for improving the system.

[0282] The present invention is a system for improving the exercise motivation of the elderly, and functions by the cooperation of a server, a terminal, and a user with each other.

[0283] The server collects exercise content from the database, performs analysis and standardization. For this, video recognition software and analysis tools are used. For example, using Python's OpenCV and TensorFlow, it is possible to analyze video materials and extract features. Based on the analyzed data, the server generates an exercise scenario specialized for the elderly and creates new media based on the generative AI model. Examples of the generative AI model used here include general algorithms that process text and image prompts.

[0284] The generated media is edited and optimized by the server. In this editing process, visual and auditory adjustments are made to improve the video quality using video editing software (e.g., Adobe Premiere Pro). Specifically, it includes brightness adjustment and voice synthesis.

[0285] The optimized media is sent to the terminal and personalized based on the user's past usage history and interests. The terminal analyzes the user's profile using a machine learning model and provides optimal content. As a result, content suitable for each user is proposed, and a high level of exercise motivation is maintained.

[0286] The user watches the provided video and performs exercises according to the instructions. After that, it is possible to input experience-based feedback to the terminal. The server collects this feedback and continues to improve the accuracy and quality of the system by utilizing it for the next media generation.

[0287] As a specific example, it is possible to generate corresponding media by inputting a prompt sentence such as "Please generate a video of content that the elderly can enjoy and enhance their exercise motivation" into the generation AI model. Through this process, viewers are able to continuously increase their motivation for exercise.

[0288] The flow of the specific process in Example 1 will be described using FIG. 11.

[0289] Step 1:

[0290] The server collects exercise content for the elderly from the database. As input, metadata and URLs of exercise videos are provided, and as output, raw data for analysis and standardization is obtained. As a specific operation, a query is sent to the database through an API, and the corresponding video file is acquired.

[0291] Step 2:

[0292] The server analyzes and standardizes the collected motion content. Raw data obtained in step 1 is passed as input, and standardized data is generated as output. For data processing, important movements within the video are extracted using Python's OpenCV, and these movements are analyzed using TensorFlow. Specifically, each video frame is analyzed, and motion features are extracted as numerical data.

[0293] Step 3:

[0294] The server generates scenarios based on standardized data. The input is the output data from step 2. The output is a scenario designed to evoke emotions in the viewer. Specifically, it uses an emotion analysis algorithm to analyze which scenes evoke which emotions and constructs the optimal scenario.

[0295] Step 4:

[0296] The server generates media using a generative AI model based on the generated scenario. The input consists of a scenario and a prompt (e.g., "Please generate a video that elderly people can enjoy and that will motivate them to exercise"), and the output is the creation of new media. Specifically, the prompt is input to the generative AI model, and the AI ​​automatically outputs media combining video clips and audio.

[0297] Step 5:

[0298] The server edits the generated media and optimizes its visual and auditory elements. The input is the output media from step 4. The output is the visually and aurally optimized media. Specifically, Adobe Premiere Pro is used to edit video transitions and background music to provide a pleasant viewing experience.

[0299] Step 6:

[0300] The terminal receives optimized media and personalizes it based on the user's past usage history and interests. As inputs, viewing history data and optimized media are provided. As an output, personalized media is generated. As a specific operation, the terminal uses a machine learning model to analyze the user's profile, select and deliver the optimal content.

[0301] Step 7:

[0302] The user watches the delivered media, performs exercises, and then inputs feedback to the terminal. As inputs, the user's experiences and feelings are provided, and as an output, feedback data is generated. As a specific operation, the user inputs their feelings and improvement points into the form provided by the terminal and sends them to the system.

[0303] Step 8:

[0304] The server analyzes the collected feedback and utilizes it for the next media generation. As an input, feedback data from the user is passed. As an output, an improved generation model is obtained. As a specific operation, a natural language processing library is used to analyze the feedback text, classify positive and negative emotions, and reflect them in the next content generation.

[0305] (Application Example 1)

[0306] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0307] There is a problem that there is a lack of a method for continuously enhancing the exercise motivation of the elderly. In the conventional exercise recommendation methods, the provision of individualized content for maintaining interest and the generation of adaptive content based on feedback are insufficient.

[0308] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0309] In this invention, the server includes means for collecting exercise media material, means for analyzing and standardizing the collected exercise media material, and means for generating media scenarios based on the standardized data material. This makes it possible to provide appropriately personalized exercise recommendation content to individual elderly people.

[0310] "Movement media material" refers to all digital content that uses movement as its theme and is provided through visual and auditory means.

[0311] "Means of analysis and standardization" refers to techniques for analyzing collected material data and organizing it into a consistent quality and format.

[0312] A "media scenario" refers to a storyline and structure designed to evoke emotions and interest in the audience.

[0313] "Means of generating information" refers to technologies for creating new visual and auditory content based on media scenarios.

[0314] "Means of optimizing visual and auditory elements" refers to technologies that adjust the sound and visuals of content to improve the user experience.

[0315] "Means of distribution to users" refers to the methods and technologies used to deliver generated viewing content to individual users in an appropriate format.

[0316] "Means of collecting and analyzing feedback" refers to technologies that collect user reactions and data and analyze them to help in creating future content.

[0317] "Means of improving information generation models" refers to technologies for improving the content generation process based on collected feedback.

[0318] "Personalized exercise recommendation content" refers to digital content about exercise that is customized based on each user's interests and activity history.

[0319] The system for implementing this invention mainly consists of a server and terminals. The server collects and analyzes exercise media material and generates media scenarios based on standardized data. Specifically, the server uses platforms such as Google Cloud AI and Amazon SageMaker to optimize visual and auditory content. This creates new information that is of interest to the user and delivers it to the terminals. The terminals are smartphones, tablets, etc., and play the role of delivering personalized content to the user. Users can view this content and provide feedback. The server collects and analyzes user feedback to improve future content generation models. This provides more personalized exercise recommendation content and helps to stimulate users' motivation to continue exercising.

[0320] For example, if a user watches content titled "Morning stretches in the park," a new scenario, such as "Stretches in a spring park with cherry blossoms," will be generated based on their reaction. An example of a prompt for the generating AI is, "Generate a morning stretching video for seniors, using a relaxed park scene."

[0321] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0322] Step 1:

[0323] The server collects sports media material from a database. This material includes video and audio files containing specific sports scenes. For input, it sends queries to the database based on pre-configured acquisition conditions and filter information, and receives all the material. The output is a list of the material stored on the server. Specifically, it uses a Python library to retrieve material from the database and converts it into a format that can be managed on the server.

[0324] Step 2:

[0325] The server analyzes and standardizes the collected motion data. During the analysis phase, it extracts information such as the type, duration, and difficulty of the motion contained in the data. The input is the data accumulated in step 1, and the output is structured metadata. Specifically, it uses a machine learning model to classify the data and extract specific features.

[0326] Step 3:

[0327] The server generates media scenarios based on standardized data. These scenarios include a narrative structure designed to engage the audience. The inputs are the standardized data obtained in step 2 and specific scenario generation rules, and the output is a completed media scenario. Specifically, it generates a storyline constructed using natural language processing.

[0328] Step 4:

[0329] The server generates new visual and auditory content based on the generated scenario. It utilizes a generative AI model to create visual materials for the target scene. The input is the scenario and template information generated in step 3, and the output is high-quality video and audio files. Specifically, prompt text is input to the AI ​​model to generate video content corresponding to the scenario.

[0330] Step 5:

[0331] The server optimizes the generated content visually and aurally. This optimization includes color correction and volume adjustment. The input is the content obtained in step 4, and the output is the visually and aurally optimized content. Specifically, video editing software is used to improve video quality through automatic correction functions.

[0332] Step 6:

[0333] The server delivers optimized content to the terminal. This delivery is done in real time and is adjusted to the individual user's preferences. The input is the content prepared in step 5, and the output is the video stream sent to the user's terminal. Specifically, data transfer is performed using streaming technology.

[0334] Step 7:

[0335] Users view the provided content and provide feedback. This feedback records their impressions after viewing the content and details of the exercises they actually performed. The input is the user's reaction to viewing the content, and the output is structured feedback data. Specifically, data is entered through the user interface using web forms and questionnaires.

[0336] Step 8:

[0337] The server analyzes the collected feedback data and uses it to improve the content it generates next. The input is the feedback data collected in step 7, and the output is the content generation model updated based on the feedback. Specifically, it uses data analysis tools to statistically analyze the feedback and uses it as training data for the AI ​​model.

[0338] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0339] This invention combines an emotional engine with a system designed to enhance exercise motivation, and is primarily aimed at the elderly. The system consists of a server and terminals and functions through a process that effectively stimulates the user's desire to exercise.

[0340] The server first collects video material from a database, analyzes it, and standardizes it. This video material includes exercises and challenges that are easy for seniors to participate in. Next, based on the collected data, the server generates scenarios that make viewers think, "I want to try that myself," and creates new videos using a video generation model.

[0341] The generated video is edited, and its visual and auditory elements are optimized. At this stage, the device delivers the video to the user and recognizes the user's emotions in real time through an emotion engine. The emotion engine analyzes the user's facial expressions and body movements to determine what emotions they are experiencing. Based on this information, the device instantly adjusts the video content to keep the user as engaged as possible.

[0342] Users move their bodies while watching the streamed videos. During this time, they can input their feelings and post-exercise feedback through their device. This feedback is sent to the server and stored in a database along with the analysis results obtained by the emotion engine. The server uses this information to improve future video generation and emotion recognition, thereby adjusting the system to enhance its overall effectiveness.

[0343] For example, if a user smiles while watching a video, the emotion engine recognizes this, and the device switches to displaying a video with more enjoyable elements. In this way, users enjoy the process of exercising, and an exercise habit is naturally formed. Furthermore, the level of exertion the user is experiencing is recorded through feedback, allowing for the optimization of content to suit individual needs.

[0344] This system, equipped with a dynamic content adjustment engine based on emotions, can strengthen the motivation of older adults to engage in exercise spontaneously and contribute to maintaining cognitive function.

[0345] The following describes the processing flow.

[0346] Step 1:

[0347] The server extracts exercise videos for seniors from its database. These videos include a wide variety of content, from simple exercises to scenes of deliberate failure.

[0348] Step 2:

[0349] The server analyzes the collected material, standardizing it by unifying the format and adjusting the resolution as needed. This standardization makes it suitable as input for the generative model.

[0350] Step 3:

[0351] The server generates scenarios based on the user's past viewing history and preferences. These scenarios effectively combine scenes of failure and scenes of potential success.

[0352] Step 4:

[0353] The server uses the generated scenario to create new video content using a video generation model. This generation process is designed to motivate the viewer.

[0354] Step 5:

[0355] The server edits the generated video content and optimizes its visual and auditory elements by adding music and text messages, thereby enhancing the emotional appeal of the video.

[0356] Step 6:

[0357] The device delivers edited video to the user. Equipped with an emotion engine, it analyzes the user's facial expressions and movements in real time and recognizes their emotions.

[0358] Step 7:

[0359] The emotion engine analyzes the user's emotions while viewing content and provides the results to the device, instantly adjusting the video content as needed. This ensures that content tailored to the user is continuously displayed.

[0360] Step 8:

[0361] Users provide their emotions and reactions to the movement to the device, which are recorded as feedback after viewing the video. This feedback is then collected.

[0362] Step 9:

[0363] The server analyzes the sentiment analysis results and user feedback sent from the terminal and uses this information to improve the next video generation process. This enhances the overall efficiency of the system.

[0364] (Example 2)

[0365] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0366] In modern society, the impact of lack of exercise on the health of older adults is serious. Traditional methods of promoting exercise are not always suitable for all older adults, and problems such as a lack of motivation and inability to adapt to individual needs have emerged. In particular, there are situations where it is difficult to effectively stimulate the desire to exercise.

[0367] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0368] In this invention, the server includes means for collecting exercise-related data, means for analyzing and standardizing the collected exercise-related data, and means for generating video scenarios based on the standardized data. This enables dynamic exercise promotion that responds to the user's interests and emotions.

[0369] "Exercise-related data" refers to information related to the exercise a user performs, including the type, intensity, and frequency of exercise.

[0370] "Analysis" refers to the process of thoroughly examining collected data and transforming its contents into an understandable format.

[0371] "Standardization" refers to the process of organizing data from different formats and standards according to a unified set of criteria.

[0372] A "video scenario" refers to a plan or storyline for video production, which comprises content that will interest the user.

[0373] "Editing" refers to the process of modifying video and audio to create a final work.

[0374] "Visual and auditory elements" refer to all the components in a video that a user can perceive visually and aurally.

[0375] "Emotional recognition" refers to technology that uses sensors to read a user's facial expressions and actions and uses that information to determine their emotional state.

[0376] "Dynamic adjustment" refers to the process of receiving information in real time and changing content or its elements in response to that information.

[0377] "Response" refers to the reaction or feedback that a user inputs into the system.

[0378] "Analyzed response" refers to information obtained as a result of the system analyzing user feedback.

[0379] "Personalization" refers to the process of tailoring the content provided to each user to their individual interests and history.

[0380] "Observation history" refers to records of what a user has watched in the past and their viewing trends.

[0381] This invention is a system designed to increase the motivation of elderly and other users to exercise. This system works in conjunction with a server and a terminal to provide users with personalized exercise videos.

[0382] The server collects exercise-related data from a database. This includes video footage of simple exercises suitable for seniors, as well as some with a slight challenge element. The server analyzes this data and standardizes and unifies the format. Video editing software is used in this process to organize the data format and make it reusable.

[0383] Next, the server uses a generative AI model to generate video scenarios that will capture the viewer's interest. This involves using prompts, such as "Create a scene of light exercise for seniors to enjoy," to input instructions into the generative AI model. Based on this scenario, the newly generated video is adjusted, and its visual and auditory elements are optimized. Specifically, CG technology is used to prepare exercise scenes suitable for the user.

[0384] The device delivers the generated video to the user. When the user watches the delivered video, the emotion engine recognizes the user's emotions in real time. This recognition uses technology that reads the user's facial expressions and body movements using cameras and sensors. For example, if the user smiles, the emotion engine will determine that the user is "enjoying" the expression, and the device will automatically adjust the video content based on that feedback.

[0385] Users move their bodies while watching these videos, and after exercising, they input feedback into their device. This feedback is sent to a server and stored in a database for analysis. The analyzed feedback information is used to improve future video generation and optimize the entire system.

[0386] This system allows users to exercise while enjoying videos tailored to their interests and physical condition, which is expected to increase their motivation to exercise and help them establish regular exercise habits.

[0387] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0388] Step 1:

[0389] The server collects exercise-related data from a database. This input data includes video materials of exercises that are easy for seniors to participate in. The server classifies these materials and selects content suitable for the target users. Specifically, videos containing simple exercises such as stretching and walking are extracted.

[0390] Step 2:

[0391] The server analyzes and standardizes the collected motion-related data. The analysis process corrects inconsistencies in video format and resolution, generating a unified format. The input to this process is the collected raw data, and the output is reusable, standardized data. Editing software is used to organize the data in different formats.

[0392] Step 3:

[0393] The server uses a generative AI model to generate video scenarios based on standardized data. During this process, the prompt "Create a scene of light exercises for seniors to enjoy" is input, and the AI ​​generates the scenario. The output is scenario data containing content that seniors would find interesting, and this is used in the next step to generate the video.

[0394] Step 4:

[0395] The server generates new video using the generated scenario data. At this stage, CG technology is used to create exercise scenes for the elderly. The input is scenario data, and the output is generated video for distribution to users. The generated video has enhanced visual and auditory elements.

[0396] Step 5:

[0397] The terminal delivers the generated video to the user. Streaming technology is used for the user's device to ensure smooth playback. The input is the generated video from the server, and the output is the user's viewing screen.

[0398] Step 6:

[0399] The emotion engine built into the device recognizes the user's emotions in real time. Cameras and sensors capture the user's facial expressions and body movements, providing them as input information. The output of the emotion analysis is data indicating the user's emotional state, which is used for video adjustment.

[0400] Step 7:

[0401] The device dynamically adjusts the video content based on the output of the emotion engine. For example, if the user smiles to indicate enjoyment, more cheerful scenes will be added to the video. The input is emotion data, and the output is the adjusted video displayed to the user.

[0402] Step 8:

[0403] Users enter feedback into a device after their workout. This feedback includes their workout experience and impressions. This feedback is sent from the device to a server and used for analysis.

[0404] Step 9:

[0405] The server integrates feedback and analysis results from the emotion engine to optimize the system. The input to this process is user feedback data, and the output is improved system specifications for future video generation and content adjustments.

[0406] (Application Example 2)

[0407] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0408] In industrial machinery, reducing the mental and physical burden on users and maintaining sustained motivation are crucial for supporting efficient work execution and improving productivity. However, conventional methods have made it difficult to provide appropriate feedback and adjustments based on the machine's condition and the user's mental state. This leads to decreased work efficiency and, consequently, economic losses, posing a significant challenge.

[0409] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0410] In this invention, the server includes means for collecting motion data material, means for analyzing and standardizing the collected motion data material, and means for generating information scenarios based on the standardized data material. This makes it possible to recognize and analyze the user's emotions in real time and dynamically adjust the machine's work content and pace accordingly.

[0411] "Motion data material" refers to data related to machine operation information and user work movements, and is used to propose efficient work processes.

[0412] "Analysis and standardization" is the process of converting collected data into a basic format and preparing the data necessary to generate new information.

[0413] An "information scenario" is a set of instructions for a series of actions or reactions, generated based on analyzed and standardized data.

[0414] "Expression" refers to content that is generated using an information scenario and presented to the user through visual and auditory elements.

[0415] "Emotional state" refers to the current psychological and emotional condition of the user, based on their facial expressions and actions.

[0416] "Dynamic adjustment" means changing the content and pace in real time according to the user's emotional state.

[0417] "Feedback" refers to responses and information provided based on the user's experiences and evaluations, and is used as material for improving the system.

[0418] To realize this invention, a server-centric system is built, starting with collecting motion data, analyzing it, and standardizing it. Based on the collected data, the server generates information scenarios and uses these scenarios to form representations. The generated representations are edited to optimize visual and auditory elements and delivered to users via terminals.

[0419] The server utilizes cameras and sensors mounted on the machine to recognize the user's emotional state in real time and analyze the user's facial expressions and movements. Specific examples include the use of emotion analysis libraries such as OpenCV and dlib. This allows for dynamic adjustment of the content of the expressions in response to the user's psychological reactions.

[0420] Users can adjust their work pace and content through the delivered expressions, thereby improving efficiency. For example, if fatigue is detected while a user is performing a certain task, the server can provide expressions intended to promote relaxation, thereby improving the user's work efficiency.

[0421] Furthermore, user feedback is sent to the server via the device. This feedback information will be used to improve expression generation and emotion recognition technologies in the future. Through this iterative learning process, the accuracy of the generative AI model will be improved.

[0422] An example of a prompt message might be, "If the robot shows signs of fatigue during work, generate a refresh video and provide a positive scenario to encourage a break." This allows the server to generate an appropriate scenario based on the instructions, creating a dynamic work environment.

[0423] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0424] Step 1:

[0425] The server collects motion data material through various sensors and cameras. This input data includes videos of work movements and environmental conditions such as temperature. The server converts this data into a digital format and stores it in a database.

[0426] Step 2:

[0427] The server analyzes and standardizes the collected motion data. It uses raw data from the database as input and performs data processing such as noise reduction and data transformation. The standardized data forms the basis for the next processing step.

[0428] Step 3:

[0429] The server generates information scenarios based on standardized data. Utilizing a generative AI model, the algorithm develops the optimal scenario based on the prompt. This output is saved as an operational program and used in the next step.

[0430] Step 4:

[0431] The server generates visual and auditory representations using the generated information scenarios. It performs data calculations to create an effective combination of video and audio, and outputs the results as representation data.

[0432] Step 5:

[0433] The server edits and optimizes the generated representation data. Data processing, such as filtering and adding effects, is performed to balance the visual and auditory elements. This optimized output is then ready to be presented to the user.

[0434] Step 6:

[0435] The terminal delivers optimized content to the user. The user views the presented content and uses it as a guide for their work. Because this delivery is done in real time, a network protocol for delivery is used.

[0436] Step 7:

[0437] The device recognizes the user's emotional state in real time using a camera and sensors. The data obtained as input is analyzed using libraries such as OpenCV to identify the emotional state. The analysis results are then sent to a server.

[0438] Step 8:

[0439] The server dynamically adjusts the content of its expressions based on sentiment analysis results. Using prompt sentences as a guide, the generative AI model updates the scenario and outputs appropriate content in real time to keep the user interested.

[0440] Step 9:

[0441] Users input subjective experiences and evaluations into their devices through feedback. This input data is sent to a server for improvement and used in future representation generation models.

[0442] Step 10:

[0443] The server uses the collected feedback and emotion recognition results to improve its expression generation model. By analyzing input data and training the generative AI model, it enhances the accuracy of future scenario generation.

[0444] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0445] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0446] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0447] [Third Embodiment]

[0448] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0449] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0450] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0451] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0452] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0453] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0454] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0455] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0456] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0457] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0458] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0459] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0460] This invention is a system for increasing exercise motivation, primarily targeting the elderly. This system effectively functions through a series of processes utilizing a server and terminals. The server first collects exercise video material for the elderly from a database, analyzes it, and standardizes it. The standardized material is selected based on scenarios designed to evoke specific emotions.

[0461] The server operates a video generation model based on these selected scenarios to generate new footage. This footage combines scenes of deliberate failure with scenes showing the possibility of success, designed to motivate viewers to take on the challenge themselves. The generated footage is edited and visually and aurally optimized. It is then delivered to the device in a personalized format, adjusted based on the user's interests and past viewing history.

[0462] The feelings and experiences that users have after watching the videos and actually engaging in the exercises are collected as feedback. The server analyzes this feedback and stores it as data to help generate the next videos. Through this analysis process, the system is continuously improved, enhancing its ability to more effectively motivate viewers to exercise.

[0463] For example, a user might watch a video streamed on their device and be inspired to try stretching themselves. This user's data is collected, and their post-viewing feedback and the extent of their exercise are analyzed. Based on this analysis, more refined content tailored to that user can be delivered in the future. This method encourages users to continue exercising enthusiastically, which in turn contributes to maintaining and improving cognitive function.

[0464] The following describes the processing flow.

[0465] Step 1:

[0466] The server collects exercise videos for seniors from a database. This material includes both fully successful and intentionally unsuccessful exercises.

[0467] Step 2:

[0468] The server analyzes the collected video footage, unifying and standardizing its resolution and format. This prepares the video generation model for smooth processing.

[0469] Step 3:

[0470] The server designs scenarios that combine failure and success scenes from video footage to effectively evoke emotions in viewers. These scenarios are built based on past viewing data.

[0471] Step 4:

[0472] The server executes a video generation model based on the selected scenario and generates new footage. This footage is structured to make viewers think, "I want to do this myself."

[0473] Step 5:

[0474] The server edits the generated video and optimizes it by adding visual and auditory elements. This includes background music and text messages to enhance the video's appeal.

[0475] Step 6:

[0476] The device delivers the completed video to the user. The video is delivered in a format that is individually tailored to the user's interests and viewing history.

[0477] Step 7:

[0478] Users watch the streamed videos and, inspired by them, try exercising themselves. Their impressions after watching the videos and the results of their exercise are collected as feedback.

[0479] Step 8:

[0480] The server analyzes user feedback and stores data to help generate future videos. This allows the system to continuously improve and provide more effective video content.

[0481] (Example 1)

[0482] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0483] Maintaining sustained exercise motivation among older adults is crucial for preserving their physical and cognitive health. However, traditional exercise programs are generally uniform and fail to adequately consider individual motivations and interests. Furthermore, there is a lack of mechanisms for effectively utilizing user feedback, making continuous improvement difficult. This leads to the challenge that older adults tend to lose interest in exercise.

[0484] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0485] In this invention, the server includes means for collecting exercise content, means for analyzing and standardizing the collected exercise content, and means for generating scenarios based on the standardized content. This makes it possible to stimulate the user's interest and emotions through personalized video and continuously increase their motivation to exercise.

[0486] "Exercise content" refers to media materials such as videos and audio used to promote physical activity for the elderly.

[0487] "Standardization" is the process of converting motion content from different formats and with different content into a consistent format, making it easier to analyze.

[0488] A "scenario" is the flow and structure of action content designed to evoke specific emotions or motivations in the viewer.

[0489] "Media" is a general term for files and streams containing visual and auditory elements that are generated for distribution to users.

[0490] A "generative model" is an algorithm or system that automatically generates new content based on specific prompts.

[0491] "Editing" is the process of modifying the visual and auditory elements of generated media in order to provide an optimal user experience.

[0492] "Personalization" refers to optimizing the media delivered to each user based on their past usage history and interests.

[0493] "Feedback" refers to the impressions, opinions, and experience-based information that users provide after viewing, and it is data that helps improve the system.

[0494] This invention is a system for improving the exercise motivation of the elderly, and it functions through the cooperation of the server, terminal, and user.

[0495] The server collects exercise content from a database, analyzes it, and standardizes it. This involves using video recognition software and analysis tools. For example, it is possible to analyze video material and extract features using Python's OpenCV or TensorFlow. Based on the analyzed data, the server generates exercise scenarios tailored to the elderly and creates new media based on a generative AI model. The generative AI model used here includes general algorithms that process text and image prompts.

[0496] The generated media is edited and optimized by the server. This editing process uses video editing software (e.g., Adobe Premiere Pro) to make visual and auditory adjustments to improve the quality of the video. Specifically, this includes adjusting brightness and compositing audio.

[0497] Optimized media is sent to the device and personalized based on the user's past usage history and interests. The device uses machine learning models to analyze the user's profile and provide optimal content. This suggests content tailored to each user, maintaining a high level of motivation for exercise.

[0498] Users watch the provided video and perform exercises according to the instructions. They can then input feedback based on their experience into their device. The server collects this feedback and uses it to continuously improve the system's accuracy and quality for future media generation.

[0499] As a concrete example, by inputting a prompt message into the AI ​​model such as, "Please generate videos that seniors can enjoy and that will increase their motivation to exercise," it is possible to generate corresponding media. Through this process, viewers can continuously increase their motivation to exercise.

[0500] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0501] Step 1:

[0502] The server collects exercise content for seniors from a database. The input is metadata and URLs of exercise videos, and the output is raw data for analysis and standardization. Specifically, it queries the database via an API and retrieves the corresponding video files.

[0503] Step 2:

[0504] The server analyzes and standardizes the collected motion content. Raw data obtained in step 1 is passed as input, and standardized data is generated as output. For data processing, important movements within the video are extracted using Python's OpenCV, and these movements are analyzed using TensorFlow. Specifically, each video frame is analyzed, and motion features are extracted as numerical data.

[0505] Step 3:

[0506] The server generates scenarios based on standardized data. The input is the output data from step 2. The output is a scenario designed to evoke emotions in the viewer. Specifically, it uses an emotion analysis algorithm to analyze which scenes evoke which emotions and constructs the optimal scenario.

[0507] Step 4:

[0508] The server generates media using a generative AI model based on the generated scenario. The input consists of a scenario and a prompt (e.g., "Please generate a video that elderly people can enjoy and that will motivate them to exercise"), and the output is the creation of new media. Specifically, the prompt is input to the generative AI model, and the AI ​​automatically outputs media combining video clips and audio.

[0509] Step 5:

[0510] The server edits the generated media and optimizes its visual and auditory elements. The input is the output media from step 4. The output is the visually and aurally optimized media. Specifically, Adobe Premiere Pro is used to edit video transitions and background music to provide a pleasant viewing experience.

[0511] Step 6:

[0512] The device receives optimized media and personalizes it based on the user's past usage history and interests. Viewing history data and optimized media are provided as input. Personalized media is generated as output. Specifically, the device uses a machine learning model to analyze the user's profile and select and deliver the most relevant content.

[0513] Step 7:

[0514] Users watch the delivered media, perform the exercise, and then input feedback into the device. The input consists of the user's experience and impressions, and the output is feedback data. Specifically, the user enters their impressions and suggestions for improvement into a form provided on the device and submits it to the system.

[0515] Step 8:

[0516] The server analyzes the collected feedback and uses it to improve future media generation. User feedback data is provided as input. An improved generative model is obtained as output. Specifically, a natural language processing library is used to analyze the feedback text, classify positive and negative emotions, and reflect these in the next content generation.

[0517] (Application Example 1)

[0518] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0519] There is a challenge in the lack of methods to sustainably increase the motivation for exercise among the elderly. Conventional methods of promoting exercise are insufficient in providing personalized content to maintain interest and in generating adaptive content based on feedback.

[0520] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0521] In this invention, the server includes means for collecting exercise media material, means for analyzing and standardizing the collected exercise media material, and means for generating media scenarios based on the standardized data material. This makes it possible to provide appropriately personalized exercise recommendation content to individual elderly people.

[0522] "Movement media material" refers to all digital content that uses movement as its theme and is provided through visual and auditory means.

[0523] "Means of analysis and standardization" refers to techniques for analyzing collected material data and organizing it into a consistent quality and format.

[0524] A "media scenario" refers to a storyline and structure designed to evoke emotions and interest in the audience.

[0525] "Means of generating information" refers to technologies for creating new visual and auditory content based on media scenarios.

[0526] "Means of optimizing visual and auditory elements" refers to technologies that adjust the sound and visuals of content to improve the user experience.

[0527] "Means of distribution to users" refers to the methods and technologies used to deliver generated viewing content to individual users in an appropriate format.

[0528] "Means of collecting and analyzing feedback" refers to technologies that collect user reactions and data and analyze them to help in creating future content.

[0529] "Means of improving information generation models" refers to technologies for improving the content generation process based on collected feedback.

[0530] "Personalized exercise recommendation content" refers to digital content about exercise that is customized based on each user's interests and activity history.

[0531] The system for implementing this invention mainly consists of a server and terminals. The server collects and analyzes exercise media material and generates media scenarios based on standardized data. Specifically, the server uses platforms such as Google Cloud AI and Amazon SageMaker to optimize visual and auditory content. This creates new information that is of interest to the user and delivers it to the terminals. The terminals are smartphones, tablets, etc., and play the role of delivering personalized content to the user. Users can view this content and provide feedback. The server collects and analyzes user feedback to improve future content generation models. This provides more personalized exercise recommendation content and helps to stimulate users' motivation to continue exercising.

[0532] For example, if a user watches content titled "Morning stretches in the park," a new scenario, such as "Stretches in a spring park with cherry blossoms," will be generated based on their reaction. An example of a prompt for the generating AI is, "Generate a morning stretching video for seniors, using a relaxed park scene."

[0533] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0534] Step 1:

[0535] The server collects sports media material from a database. This material includes video and audio files containing specific sports scenes. For input, it sends queries to the database based on pre-configured acquisition conditions and filter information, and receives all the material. The output is a list of the material stored on the server. Specifically, it uses a Python library to retrieve material from the database and converts it into a format that can be managed on the server.

[0536] Step 2:

[0537] The server analyzes and standardizes the collected motion data. During the analysis phase, it extracts information such as the type, duration, and difficulty of the motion contained in the data. The input is the data accumulated in step 1, and the output is structured metadata. Specifically, it uses a machine learning model to classify the data and extract specific features.

[0538] Step 3:

[0539] The server generates media scenarios based on standardized data. These scenarios include a narrative structure designed to engage the audience. The inputs are the standardized data obtained in step 2 and specific scenario generation rules, and the output is a completed media scenario. Specifically, it generates a storyline constructed using natural language processing.

[0540] Step 4:

[0541] The server generates new visual and auditory content based on the generated scenario. It utilizes a generative AI model to create visual materials for the target scene. The input is the scenario and template information generated in step 3, and the output is high-quality video and audio files. Specifically, prompt text is input to the AI ​​model to generate video content corresponding to the scenario.

[0542] Step 5:

[0543] The server optimizes the generated content visually and aurally. This optimization includes color correction and volume adjustment. The input is the content obtained in step 4, and the output is the visually and aurally optimized content. Specifically, video editing software is used to improve video quality through automatic correction functions.

[0544] Step 6:

[0545] The server delivers optimized content to the terminal. This delivery is done in real time and is adjusted to the individual user's preferences. The input is the content prepared in step 5, and the output is the video stream sent to the user's terminal. Specifically, data transfer is performed using streaming technology.

[0546] Step 7:

[0547] Users view the provided content and provide feedback. This feedback records their impressions after viewing the content and details of the exercises they actually performed. The input is the user's reaction to viewing the content, and the output is structured feedback data. Specifically, data is entered through the user interface using web forms and questionnaires.

[0548] Step 8:

[0549] The server analyzes the collected feedback data and uses it to improve the content it generates next. The input is the feedback data collected in step 7, and the output is the content generation model updated based on the feedback. Specifically, it uses data analysis tools to statistically analyze the feedback and uses it as training data for the AI ​​model.

[0550] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0551] This invention combines an emotional engine with a system designed to enhance exercise motivation, and is primarily aimed at the elderly. The system consists of a server and terminals and functions through a process that effectively stimulates the user's desire to exercise.

[0552] The server first collects video material from a database, analyzes it, and standardizes it. This video material includes exercises and challenges that are easy for seniors to participate in. Next, based on the collected data, the server generates scenarios that make viewers think, "I want to try that myself," and creates new videos using a video generation model.

[0553] The generated video is edited, and its visual and auditory elements are optimized. At this stage, the device delivers the video to the user and recognizes the user's emotions in real time through an emotion engine. The emotion engine analyzes the user's facial expressions and body movements to determine what emotions they are experiencing. Based on this information, the device instantly adjusts the video content to keep the user as engaged as possible.

[0554] Users move their bodies while watching the streamed videos. During this time, they can input their feelings and post-exercise feedback through their device. This feedback is sent to the server and stored in a database along with the analysis results obtained by the emotion engine. The server uses this information to improve future video generation and emotion recognition, thereby adjusting the system to enhance its overall effectiveness.

[0555] For example, if a user smiles while watching a video, the emotion engine recognizes this, and the device switches to displaying a video with more enjoyable elements. In this way, users enjoy the process of exercising, and an exercise habit is naturally formed. Furthermore, the level of exertion the user is experiencing is recorded through feedback, allowing for the optimization of content to suit individual needs.

[0556] This system, equipped with a dynamic content adjustment engine based on emotions, can strengthen the motivation of older adults to engage in exercise spontaneously and contribute to maintaining cognitive function.

[0557] The following describes the processing flow.

[0558] Step 1:

[0559] The server extracts exercise videos for seniors from its database. These videos include a wide variety of content, from simple exercises to scenes of deliberate failure.

[0560] Step 2:

[0561] The server analyzes the collected material, standardizing it by unifying the format and adjusting the resolution as needed. This standardization makes it suitable as input for the generative model.

[0562] Step 3:

[0563] The server generates scenarios based on the user's past viewing history and preferences. These scenarios effectively combine scenes of failure and scenes of potential success.

[0564] Step 4:

[0565] The server uses the generated scenario to create new video content using a video generation model. This generation process is designed to be engaging and motivating for the viewer.

[0566] Step 5:

[0567] The server edits the generated video content and optimizes its visual and auditory elements by adding music and text messages, thereby enhancing the emotional appeal of the video.

[0568] Step 6:

[0569] The device delivers edited video to the user. Equipped with an emotion engine, it analyzes the user's facial expressions and movements in real time and recognizes their emotions.

[0570] Step 7:

[0571] The emotion engine analyzes the user's emotions while viewing content and provides the results to the device, instantly adjusting the video content as needed. This ensures that content tailored to the user is continuously displayed.

[0572] Step 8:

[0573] Users provide their emotions and reactions to the movement to the device, which are recorded as feedback after viewing the video. This feedback is then collected.

[0574] Step 9:

[0575] The server analyzes the sentiment analysis results and user feedback sent from the terminal and uses this information to improve the next video generation process. This enhances the overall efficiency of the system.

[0576] (Example 2)

[0577] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0578] In modern society, the impact of lack of exercise on the health of older adults is serious. Traditional methods of promoting exercise are not always suitable for all older adults, and problems such as a lack of motivation and inability to adapt to individual needs have emerged. In particular, there are situations where it is difficult to effectively stimulate the desire to exercise.

[0579] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0580] In this invention, the server includes means for collecting exercise-related data, means for analyzing and standardizing the collected exercise-related data, and means for generating video scenarios based on the standardized data. This enables dynamic exercise promotion that responds to the user's interests and emotions.

[0581] "Exercise-related data" refers to information related to the exercise a user performs, including the type, intensity, and frequency of exercise.

[0582] "Analysis" refers to the process of thoroughly examining collected data and transforming its contents into an understandable format.

[0583] "Standardization" refers to the process of organizing data from different formats and standards according to a unified set of criteria.

[0584] A "video scenario" refers to a plan or storyline for video production, which comprises content that will interest the user.

[0585] "Editing" refers to the process of modifying video and audio to create a final work.

[0586] "Visual and auditory elements" refer to all the components in a video that a user can perceive visually and aurally.

[0587] "Emotional recognition" refers to technology that uses sensors to read a user's facial expressions and actions and uses that information to determine their emotional state.

[0588] "Dynamic adjustment" refers to the process of receiving information in real time and changing content or its elements in response to that information.

[0589] "Response" refers to the reaction or feedback that a user inputs into the system.

[0590] "Analyzed response" refers to information obtained as a result of the system analyzing user feedback.

[0591] "Personalization" refers to the process of tailoring the content provided to each user to their individual interests and history.

[0592] "Observation history" refers to records of what a user has watched in the past and their viewing trends.

[0593] This invention is a system designed to increase the motivation of elderly and other users to exercise. This system operates through the coordinated action of a server and a terminal, enabling the provision of personalized exercise videos to the user.

[0594] The server collects exercise-related data from a database. This includes video footage of simple exercises suitable for seniors, as well as some with a slight challenge element. The server analyzes this data and standardizes and unifies the format. Video editing software is used in this process to organize the data format and make it reusable.

[0595] Next, the server uses a generative AI model to generate video scenarios that will capture the viewer's interest. This involves using prompts, such as "Create a scene of light exercise for seniors to enjoy," to input instructions into the generative AI model. Based on this scenario, the newly generated video is adjusted, and its visual and auditory elements are optimized. Specifically, CG technology is used to prepare exercise scenes suitable for the user.

[0596] The device delivers the generated video to the user. When the user watches the delivered video, the emotion engine recognizes the user's emotions in real time. This recognition uses technology that reads the user's facial expressions and body movements using cameras and sensors. For example, if the user smiles, the emotion engine will determine that the user is "enjoying" the expression, and the device will automatically adjust the video content based on that feedback.

[0597] Users move their bodies while watching these videos, and after exercising, they input feedback into their device. This feedback is sent to a server and stored in a database for analysis. The analyzed feedback information is used to optimize future video generation and the overall system.

[0598] This system allows users to exercise while enjoying videos tailored to their interests and physical condition, which is expected to increase their motivation to exercise and help them establish regular exercise habits.

[0599] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0600] Step 1:

[0601] The server collects exercise-related data from a database. This input data includes video materials of exercises that are easy for seniors to participate in. The server classifies these materials and selects content suitable for the target users. Specifically, videos containing simple exercises such as stretching and walking are extracted.

[0602] Step 2:

[0603] The server analyzes and standardizes the collected motion-related data. The analysis process corrects inconsistencies in video format and resolution, generating a unified format. The input to this process is the collected raw data, and the output is reusable, standardized data. Editing software is used to organize the data in different formats.

[0604] Step 3:

[0605] The server uses a generative AI model to generate video scenarios based on standardized data. During this process, the prompt "Create a scene of light exercises for seniors to enjoy" is input, and the AI ​​generates the scenario. The output is scenario data containing content that seniors would find interesting, and this is used in the next step to generate the video.

[0606] Step 4:

[0607] The server generates new video using the generated scenario data. At this stage, CG technology is used to create exercise scenes for the elderly. The input is scenario data, and the output is generated video for distribution to users. The generated video has enhanced visual and auditory elements.

[0608] Step 5:

[0609] The terminal delivers the generated video to the user. Streaming technology is used for the user's device to ensure smooth playback. The input is the generated video from the server, and the output is the user's viewing screen.

[0610] Step 6:

[0611] The emotion engine built into the device recognizes the user's emotions in real time. Cameras and sensors capture the user's facial expressions and body movements, providing them as input information. The output of the emotion analysis is data indicating the user's emotional state, and is used for video adjustment.

[0612] Step 7:

[0613] The device dynamically adjusts the video content based on the output of the emotion engine. For example, if the user smiles to indicate enjoyment, more cheerful scenes will be added to the video. The input is emotion data, and the output is the adjusted video displayed to the user.

[0614] Step 8:

[0615] Users enter feedback into a device after their workout. This feedback includes their workout experience and impressions. This feedback is sent from the device to a server and used for analysis.

[0616] Step 9:

[0617] The server integrates feedback and analysis results from the emotion engine to optimize the system. The input to this process is user feedback data, and the output is improved system specifications for future video generation and content adjustments.

[0618] (Application Example 2)

[0619] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0620] In industrial machinery, reducing the mental and physical burden on users and maintaining sustained motivation are crucial for supporting efficient work execution and improving productivity. However, conventional methods have made it difficult to provide appropriate feedback and adjustments based on the machine's condition and the user's mental state. This leads to decreased work efficiency and, consequently, economic losses, posing a significant challenge.

[0621] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0622] In this invention, the server includes means for collecting motion data material, means for analyzing and standardizing the collected motion data material, and means for generating information scenarios based on the standardized data material. This makes it possible to recognize and analyze the user's emotions in real time and dynamically adjust the machine's work content and pace accordingly.

[0623] "Motion data material" refers to data related to machine operation information and user work movements, and is used to propose efficient work processes.

[0624] "Analysis and standardization" is the process of converting collected data into a basic format and preparing the data necessary to generate new information.

[0625] An "information scenario" is a set of instructions for a series of actions or reactions, generated based on analyzed and standardized data.

[0626] "Expression" refers to content that is generated using an information scenario and presented to the user through visual and auditory elements.

[0627] "Emotional state" refers to the current psychological and emotional condition of the user, based on their facial expressions and actions.

[0628] "Dynamic adjustment" means changing the content and pace in real time according to the user's emotional state.

[0629] "Feedback" refers to responses and information provided based on the user's experiences and evaluations, and is used as material for improving the system.

[0630] To realize this invention, a server-centric system is built, starting with collecting motion data, analyzing it, and standardizing it. Based on the collected data, the server generates information scenarios and uses these scenarios to form representations. The generated representations are edited to optimize visual and auditory elements and delivered to users via terminals.

[0631] The server utilizes cameras and sensors mounted on the machine to recognize the user's emotional state in real time and analyze the user's facial expressions and movements. Specific examples include the use of emotion analysis libraries such as OpenCV and dlib. This allows for dynamic adjustment of the content of the expressions in response to the user's psychological reactions.

[0632] Users can adjust their work pace and content through the delivered expressions, thereby improving efficiency. For example, if fatigue is detected while a user is performing a certain task, the server can provide expressions intended to promote relaxation, thereby improving the user's work efficiency.

[0633] Furthermore, user feedback is sent to the server via the device. This feedback information will be used to improve expression generation and emotion recognition technologies in the future. Through this iterative learning process, the accuracy of the generative AI model will be improved.

[0634] An example of a prompt message might be, "If the robot shows signs of fatigue during work, generate a refresh video and provide a positive scenario to encourage a break." This allows the server to generate an appropriate scenario based on the instructions, creating a dynamic work environment.

[0635] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0636] Step 1:

[0637] The server collects motion data material through various sensors and cameras. This input data includes videos of work movements and environmental conditions such as temperature. The server converts this data into a digital format and stores it in a database.

[0638] Step 2:

[0639] The server analyzes and standardizes the collected motion data. It uses raw data from the database as input and performs data processing such as noise reduction and data transformation. The standardized data forms the basis for the next processing step.

[0640] Step 3:

[0641] The server generates information scenarios based on standardized data. Utilizing a generative AI model, the algorithm develops the optimal scenario based on the prompt. This output is saved as an operational program and used in the next step.

[0642] Step 4:

[0643] The server generates visual and auditory representations using the generated information scenarios. It performs data calculations to create an effective combination of video and audio, and outputs the results as representation data.

[0644] Step 5:

[0645] The server edits and optimizes the generated representation data. Data processing, such as filtering and adding effects, is performed to balance the visual and auditory elements. This optimized output is then ready to be presented to the user.

[0646] Step 6:

[0647] The terminal delivers optimized content to the user. The user views the presented content and uses it as a guide for their work. Because this delivery is done in real time, a network protocol for delivery is used.

[0648] Step 7:

[0649] The device uses a camera and sensors to recognize the user's emotional state in real time. The data obtained as input is analyzed using libraries such as OpenCV to identify the emotional state. The analysis results are then sent to a server.

[0650] Step 8:

[0651] The server dynamically adjusts the content of its expressions based on sentiment analysis results. Using prompt sentences as a guide, the generative AI model updates the scenario and outputs appropriate content in real time to keep the user interested.

[0652] Step 9:

[0653] Users input subjective experiences and evaluations into their devices through feedback. This input data is sent to a server for improvement and used in future representation generation models.

[0654] Step 10:

[0655] The server uses the collected feedback and emotion recognition results to improve its expression generation model. By analyzing input data and training the generative AI model, it enhances the accuracy of future scenario generation.

[0656] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0657] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0658] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0659] [Fourth Embodiment]

[0660] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0661] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0662] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0663] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0664] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0665] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0666] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0667] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0668] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0669] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0670] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0671] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0672] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0673] This invention is a system for increasing exercise motivation, primarily targeting the elderly. This system effectively functions through a series of processes utilizing a server and terminals. The server first collects exercise video material for the elderly from a database, analyzes it, and standardizes it. The standardized material is selected based on scenarios designed to evoke specific emotions.

[0674] The server operates a video generation model based on these selected scenarios to generate new footage. This footage combines scenes of deliberate failure with scenes showing the possibility of success, designed to motivate viewers to take on the challenge themselves. The generated footage is edited and visually and aurally optimized. It is then delivered to the device in a personalized format, adjusted based on the user's interests and past viewing history.

[0675] The feelings and experiences that users have after watching the videos and actually engaging in the exercises are collected as feedback. The server analyzes this feedback and stores it as data to help generate the next videos. Through this analysis process, the system is continuously improved, enhancing its ability to more effectively motivate viewers to exercise.

[0676] For example, a user might watch a video streamed on their device and be inspired to try stretching themselves. This user's data is collected, and their post-viewing feedback and the extent of their exercise are analyzed. Based on this analysis, more refined content tailored to that user can be delivered in the future. This method encourages users to continue exercising enthusiastically, which in turn contributes to maintaining and improving cognitive function.

[0677] The following describes the processing flow.

[0678] Step 1:

[0679] The server collects exercise videos for seniors from a database. This material includes both fully successful and intentionally unsuccessful exercises.

[0680] Step 2:

[0681] The server analyzes the collected video footage, unifying and standardizing its resolution and format. This prepares the video generation model for smooth processing.

[0682] Step 3:

[0683] The server designs scenarios that combine failure and success scenes from video footage to effectively evoke emotions in viewers. These scenarios are built based on past viewing data.

[0684] Step 4:

[0685] The server executes a video generation model based on the selected scenario and generates new footage. This footage is structured to make viewers think, "I want to do this myself."

[0686] Step 5:

[0687] The server edits the generated video and optimizes it by adding visual and auditory elements. This includes background music and text messages to enhance the video's appeal.

[0688] Step 6:

[0689] The device delivers the completed video to the user. The video is delivered in a format that is individually tailored to the user's interests and viewing history.

[0690] Step 7:

[0691] Users watch the streamed videos and, inspired by them, try exercising themselves. Their impressions after watching the videos and the results of their exercise are collected as feedback.

[0692] Step 8:

[0693] The server analyzes user feedback and stores data to help generate future videos. This allows the system to continuously improve and provide more effective video content.

[0694] (Example 1)

[0695] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0696] Maintaining sustained exercise motivation among older adults is crucial for preserving their physical and cognitive health. However, traditional exercise programs are generally uniform and fail to adequately consider individual motivations and interests. Furthermore, there is a lack of mechanisms for effectively utilizing user feedback, making continuous improvement difficult. This leads to the challenge that older adults tend to lose interest in exercise.

[0697] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0698] In this invention, the server includes means for collecting exercise content, means for analyzing and standardizing the collected exercise content, and means for generating scenarios based on the standardized content. This makes it possible to stimulate the user's interest and emotions through personalized video and continuously increase their motivation to exercise.

[0699] "Exercise content" refers to media materials such as videos and audio used to promote physical activity for the elderly.

[0700] "Standardization" is the process of converting motion content from different formats and with different content into a consistent format, making it easier to analyze.

[0701] A "scenario" is the flow and structure of action content designed to evoke specific emotions or motivations in the viewer.

[0702] "Media" is a general term for files and streams containing visual and auditory elements that are generated for distribution to users.

[0703] A "generative model" is an algorithm or system that automatically generates new content based on specific prompts.

[0704] "Editing" is the process of modifying the visual and auditory elements of generated media in order to provide an optimal user experience.

[0705] "Personalization" refers to optimizing the media delivered to each user based on their past usage history and interests.

[0706] "Feedback" refers to the impressions, opinions, and experience-based information that users provide after viewing, and it is data that helps improve the system.

[0707] This invention is a system for improving the exercise motivation of the elderly, and it functions through the cooperation of the server, terminal, and user.

[0708] The server collects exercise content from a database, analyzes it, and standardizes it. This involves using video recognition software and analysis tools. For example, it is possible to analyze video material and extract features using Python's OpenCV or TensorFlow. Based on the analyzed data, the server generates exercise scenarios tailored to the elderly and creates new media based on a generative AI model. The generative AI model used here includes general algorithms that process text and image prompts.

[0709] The generated media is edited and optimized by the server. This editing process uses video editing software (e.g., Adobe Premiere Pro) to make visual and auditory adjustments to improve the quality of the video. Specifically, this includes adjusting brightness and compositing audio.

[0710] Optimized media is sent to the device and personalized based on the user's past usage history and interests. The device uses machine learning models to analyze the user's profile and provide optimal content. This suggests content tailored to each user, maintaining a high level of motivation for exercise.

[0711] Users watch the provided video and perform exercises according to the instructions. They can then input feedback based on their experience into their device. The server collects this feedback and uses it to continuously improve the system's accuracy and quality for future media generation.

[0712] As a concrete example, by inputting a prompt message into the AI ​​model such as, "Please generate videos that seniors can enjoy and that will increase their motivation to exercise," it is possible to generate corresponding media. Through this process, viewers can continuously increase their motivation to exercise.

[0713] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0714] Step 1:

[0715] The server collects exercise content for seniors from a database. The input is metadata and URLs of exercise videos, and the output is raw data for analysis and standardization. Specifically, it queries the database via an API and retrieves the corresponding video files.

[0716] Step 2:

[0717] The server analyzes and standardizes the collected motion content. Raw data obtained in step 1 is passed as input, and standardized data is generated as output. For data processing, important movements within the video are extracted using Python's OpenCV, and these movements are analyzed using TensorFlow. Specifically, each video frame is analyzed, and motion features are extracted as numerical data.

[0718] Step 3:

[0719] The server generates scenarios based on standardized data. The input is the output data from step 2. The output is a scenario designed to evoke emotions in the viewer. Specifically, it uses an emotion analysis algorithm to analyze which scenes evoke which emotions and constructs the optimal scenario.

[0720] Step 4:

[0721] The server generates media using a generative AI model based on the generated scenario. The input consists of a scenario and a prompt (e.g., "Please generate a video that elderly people can enjoy and that will motivate them to exercise"), and the output is the creation of new media. Specifically, the prompt is input to the generative AI model, and the AI ​​automatically outputs media combining video clips and audio.

[0722] Step 5:

[0723] The server edits the generated media and optimizes its visual and auditory elements. The input is the output media from step 4. The output is the visually and aurally optimized media. Specifically, Adobe Premiere Pro is used to edit video transitions and background music to provide a pleasant viewing experience.

[0724] Step 6:

[0725] The device receives optimized media and personalizes it based on the user's past usage history and interests. Viewing history data and optimized media are provided as input. Personalized media is generated as output. Specifically, the device uses a machine learning model to analyze the user's profile and select and deliver the most relevant content.

[0726] Step 7:

[0727] Users watch the delivered media, perform the exercise, and then input feedback into the device. The input consists of the user's experience and impressions, and the output is feedback data. Specifically, the user enters their impressions and suggestions for improvement into a form provided on the device and submits it to the system.

[0728] Step 8:

[0729] The server analyzes the collected feedback and uses it to improve future media generation. User feedback data is provided as input. An improved generative model is obtained as output. Specifically, a natural language processing library is used to analyze the feedback text, classify positive and negative emotions, and reflect these in the next content generation.

[0730] (Application Example 1)

[0731] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0732] There is a challenge in the lack of methods to sustainably increase the motivation for exercise among the elderly. Conventional methods of promoting exercise are insufficient in providing personalized content to maintain interest and in generating adaptive content based on feedback.

[0733] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0734] In this invention, the server includes means for collecting exercise media material, means for analyzing and standardizing the collected exercise media material, and means for generating media scenarios based on the standardized data material. This makes it possible to provide appropriately personalized exercise recommendation content to individual elderly people.

[0735] "Movement media material" refers to all digital content that uses movement as its theme and is provided through visual and auditory means.

[0736] "Means of analysis and standardization" refers to techniques for analyzing collected material data and organizing it into a consistent quality and format.

[0737] A "media scenario" refers to a storyline and structure designed to evoke emotions and interest in the audience.

[0738] "Means of generating information" refers to technologies for creating new visual and auditory content based on media scenarios.

[0739] "Means of optimizing visual and auditory elements" refers to technologies that adjust the sound and visuals of content to improve the user experience.

[0740] "Means of distribution to users" refers to the methods and technologies used to deliver generated viewing content to individual users in an appropriate format.

[0741] "Means of collecting and analyzing feedback" refers to technologies that collect user reactions and data and analyze them to help in creating future content.

[0742] "Means of improving information generation models" refers to technologies for improving the content generation process based on collected feedback.

[0743] "Personalized exercise recommendation content" refers to digital content about exercise that is customized based on each user's interests and activity history.

[0744] The system for implementing this invention mainly consists of a server and terminals. The server collects and analyzes exercise media material and generates media scenarios based on standardized data. Specifically, the server uses platforms such as Google Cloud AI and Amazon SageMaker to optimize visual and auditory content. This creates new information that is of interest to the user and delivers it to the terminals. The terminals are smartphones, tablets, etc., and play the role of delivering personalized content to the user. Users can view this content and provide feedback. The server collects and analyzes user feedback to improve future content generation models. This provides more personalized exercise recommendation content and helps to stimulate users' motivation to continue exercising.

[0745] For example, if a user watches content titled "Morning stretches in the park," a new scenario, such as "Stretches in a spring park with cherry blossoms," will be generated based on their reaction. An example of a prompt for the generating AI is, "Generate a morning stretching video for seniors, using a relaxed park scene."

[0746] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0747] Step 1:

[0748] The server collects sports media material from a database. This material includes video and audio files containing specific sports scenes. For input, it sends queries to the database based on pre-configured acquisition conditions and filter information, and receives all the material. The output is a list of the material stored on the server. Specifically, it uses a Python library to retrieve material from the database and converts it into a format that can be managed on the server.

[0749] Step 2:

[0750] The server analyzes and standardizes the collected motion data. During the analysis phase, it extracts information such as the type, duration, and difficulty of the motion contained in the data. The input is the data accumulated in step 1, and the output is structured metadata. Specifically, it uses a machine learning model to classify the data and extract specific features.

[0751] Step 3:

[0752] The server generates media scenarios based on standardized data. These scenarios include a narrative structure designed to engage the audience. The inputs are the standardized data obtained in step 2 and specific scenario generation rules, and the output is a completed media scenario. Specifically, it generates a storyline constructed using natural language processing.

[0753] Step 4:

[0754] The server generates new visual and auditory content based on the generated scenario. It utilizes a generative AI model to create visual materials for the target scene. The input is the scenario and template information generated in step 3, and the output is high-quality video and audio files. Specifically, prompt text is input to the AI ​​model to generate video content corresponding to the scenario.

[0755] Step 5:

[0756] The server optimizes the generated content visually and aurally. This optimization includes color correction and volume adjustment. The input is the content obtained in step 4, and the output is the visually and aurally optimized content. Specifically, video editing software is used to improve video quality through automatic correction functions.

[0757] Step 6:

[0758] The server delivers optimized content to the terminal. This delivery is done in real time and is adjusted to the individual user's preferences. The input is the content prepared in step 5, and the output is the video stream sent to the user's terminal. Specifically, data transfer is performed using streaming technology.

[0759] Step 7:

[0760] Users view the provided content and provide feedback. This feedback records their impressions after viewing the content and details of the exercises they actually performed. The input is the user's reaction to viewing the content, and the output is structured feedback data. Specifically, data is entered through the user interface using web forms and questionnaires.

[0761] Step 8:

[0762] The server analyzes the collected feedback data and uses it to improve the content it generates next. The input is the feedback data collected in step 7, and the output is the content generation model updated based on the feedback. Specifically, it uses data analysis tools to statistically analyze the feedback and uses it as training data for the AI ​​model.

[0763] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0764] This invention combines an emotional engine with a system designed to enhance exercise motivation, and is primarily aimed at the elderly. The system consists of a server and terminals and functions through a process that effectively stimulates the user's desire to exercise.

[0765] The server first collects video material from a database, analyzes it, and standardizes it. This video material includes exercises and challenges that are easy for seniors to participate in. Next, based on the collected data, the server generates scenarios that make viewers think, "I want to try that myself," and creates new videos using a video generation model.

[0766] The generated video is edited, and its visual and auditory elements are optimized. At this stage, the device delivers the video to the user and recognizes the user's emotions in real time through an emotion engine. The emotion engine analyzes the user's facial expressions and body movements to determine what emotions they are experiencing. Based on this information, the device instantly adjusts the video content to keep the user as engaged as possible.

[0767] Users move their bodies while watching the streamed videos. During this time, they can input their feelings and post-exercise feedback through their device. This feedback is sent to the server and stored in a database along with the analysis results obtained by the emotion engine. The server uses this information to improve future video generation and emotion recognition, thereby adjusting the system to enhance its overall effectiveness.

[0768] For example, if a user smiles while watching a video, the emotion engine recognizes this, and the device switches to displaying a video with more enjoyable elements. In this way, users enjoy the process of exercising, and an exercise habit is naturally formed. Furthermore, the level of exertion the user is experiencing is recorded through feedback, allowing for the optimization of content to suit individual needs.

[0769] This system, equipped with a dynamic content adjustment engine based on emotions, can strengthen the motivation of older adults to engage in exercise spontaneously and contribute to maintaining cognitive function.

[0770] The following describes the processing flow.

[0771] Step 1:

[0772] The server extracts exercise videos for seniors from its database. These videos include a wide variety of content, from simple exercises to scenes of deliberate failure.

[0773] Step 2:

[0774] The server analyzes the collected material, standardizing it by unifying the format and adjusting the resolution as needed. This standardization makes it suitable as input for the generative model.

[0775] Step 3:

[0776] The server generates scenarios based on the user's past viewing history and preferences. These scenarios effectively combine scenes of failure and scenes of potential success.

[0777] Step 4:

[0778] The server uses the generated scenario to create new video content using a video generation model. This generation process is designed to be engaging and motivating for the viewer.

[0779] Step 5:

[0780] The server edits the generated video content and optimizes its visual and auditory elements by adding music and text messages, thereby enhancing the emotional appeal of the video.

[0781] Step 6:

[0782] The device delivers edited video to the user. Equipped with an emotion engine, it analyzes the user's facial expressions and movements in real time and recognizes their emotions.

[0783] Step 7:

[0784] The emotion engine analyzes the user's emotions while viewing content and provides the results to the device, instantly adjusting the video content as needed. This ensures that content tailored to the user is continuously displayed.

[0785] Step 8:

[0786] Users provide their emotions and reactions to the movement to the device, which are recorded as feedback after viewing the video. This feedback is then collected.

[0787] Step 9:

[0788] The server analyzes the sentiment analysis results and user feedback sent from the terminal and uses this information to improve the next video generation process. This enhances the overall efficiency of the system.

[0789] (Example 2)

[0790] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0791] In modern society, the impact of lack of exercise on the health of older adults is serious. Traditional methods of promoting exercise are not always suitable for all older adults, and problems such as a lack of motivation and inability to adapt to individual needs have emerged. In particular, there are situations where it is difficult to effectively stimulate the desire to exercise.

[0792] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0793] In this invention, the server includes means for collecting exercise-related data, means for analyzing and standardizing the collected exercise-related data, and means for generating video scenarios based on the standardized data. This enables dynamic exercise promotion that responds to the user's interests and emotions.

[0794] "Exercise-related data" refers to information related to the exercise a user performs, including the type, intensity, and frequency of exercise.

[0795] "Analysis" refers to the process of thoroughly examining collected data and transforming its contents into an understandable format.

[0796] "Standardization" refers to the process of organizing data from different formats and standards according to a unified set of criteria.

[0797] A "video scenario" refers to a plan or storyline for video production, which comprises content that will interest the user.

[0798] "Editing" refers to the process of modifying video and audio to create a final work.

[0799] "Visual and auditory elements" refer to all the components in a video that a user can perceive visually and aurally.

[0800] "Emotional recognition" refers to technology that uses sensors to read a user's facial expressions and actions and uses that information to determine their emotional state.

[0801] "Dynamic adjustment" refers to the process of receiving information in real time and changing content or its elements in response to that information.

[0802] "Response" refers to the reaction or feedback that a user inputs into the system.

[0803] "Analyzed response" refers to information obtained as a result of the system analyzing user feedback.

[0804] "Personalization" refers to the process of tailoring the content provided to each user to their individual interests and history.

[0805] "Observation history" refers to records of what a user has watched in the past and their viewing trends.

[0806] This invention is a system designed to increase the motivation of elderly and other users to exercise. This system operates through the coordinated action of a server and a terminal, enabling the provision of personalized exercise videos to the user.

[0807] The server collects exercise-related data from a database. This includes video footage of simple exercises suitable for seniors, as well as some with a slight challenge element. The server analyzes this data and standardizes and unifies the format. Video editing software is used in this process to organize the data format and make it reusable.

[0808] Next, the server uses a generative AI model to generate video scenarios that will capture the viewer's interest. This involves using prompts, such as "Create a scene of light exercise for seniors to enjoy," to input instructions into the generative AI model. Based on this scenario, the newly generated video is adjusted, and its visual and auditory elements are optimized. Specifically, CG technology is used to prepare exercise scenes suitable for the user.

[0809] The device delivers the generated video to the user. When the user watches the delivered video, the emotion engine recognizes the user's emotions in real time. This recognition uses technology that reads the user's facial expressions and body movements using cameras and sensors. For example, if the user smiles, the emotion engine will determine that the user is "enjoying" the expression, and the device will automatically adjust the video content based on that feedback.

[0810] Users move their bodies while watching these videos, and after exercising, they input feedback into their device. This feedback is sent to a server and stored in a database for analysis. The analyzed feedback information is used to optimize future video generation and the overall system.

[0811] This system allows users to exercise while enjoying videos tailored to their interests and physical condition, which is expected to increase their motivation to exercise and help them establish regular exercise habits.

[0812] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0813] Step 1:

[0814] The server collects exercise-related data from a database. This input data includes video materials of exercises that are easy for seniors to participate in. The server classifies these materials and selects content suitable for the target users. Specifically, videos containing simple exercises such as stretching and walking are extracted.

[0815] Step 2:

[0816] The server analyzes and standardizes the collected motion-related data. The analysis process corrects inconsistencies in video format and resolution, generating a unified format. The input to this process is the collected raw data, and the output is reusable, standardized data. Editing software is used to organize the data in different formats.

[0817] Step 3:

[0818] The server uses a generative AI model to generate video scenarios based on standardized data. During this process, the prompt "Create a scene of light exercises for seniors to enjoy" is input, and the AI ​​generates the scenario. The output is scenario data containing content that seniors would find interesting, and this is used in the next step to generate the video.

[0819] Step 4:

[0820] The server generates new video using the generated scenario data. At this stage, CG technology is used to create exercise scenes for the elderly. The input is scenario data, and the output is generated video for distribution to users. The generated video has enhanced visual and auditory elements.

[0821] Step 5:

[0822] The terminal delivers the generated video to the user. Streaming technology is used for the user's device to ensure smooth playback. The input is the generated video from the server, and the output is the user's viewing screen.

[0823] Step 6:

[0824] The emotion engine built into the device recognizes the user's emotions in real time. Cameras and sensors capture the user's facial expressions and body movements, providing them as input information. The output of the emotion analysis is data indicating the user's emotional state, and is used for video adjustment.

[0825] Step 7:

[0826] The device dynamically adjusts the video content based on the output of the emotion engine. For example, if the user smiles to indicate enjoyment, more cheerful scenes will be added to the video. The input is emotion data, and the output is the adjusted video displayed to the user.

[0827] Step 8:

[0828] Users enter feedback into a device after their workout. This feedback includes their workout experience and impressions. This feedback is sent from the device to a server and used for analysis.

[0829] Step 9:

[0830] The server integrates feedback and analysis results from the emotion engine to optimize the system. The input to this process is user feedback data, and the output is improved system specifications for future video generation and content adjustments.

[0831] (Application Example 2)

[0832] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0833] In industrial machinery, reducing the mental and physical burden on users and maintaining sustained motivation are crucial for supporting efficient work execution and improving productivity. However, conventional methods have made it difficult to provide appropriate feedback and adjustments based on the machine's condition and the user's mental state. This leads to decreased work efficiency and, consequently, economic losses, posing a significant challenge.

[0834] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0835] In this invention, the server includes means for collecting motion data material, means for analyzing and standardizing the collected motion data material, and means for generating information scenarios based on the standardized data material. This makes it possible to recognize and analyze the user's emotions in real time and dynamically adjust the machine's work content and pace accordingly.

[0836] "Motion data material" refers to data related to machine operation information and user work movements, and is used to propose efficient work processes.

[0837] "Analysis and standardization" is the process of converting collected data into a basic format and preparing the data necessary to generate new information.

[0838] An "information scenario" is a set of instructions for a series of actions or reactions, generated based on analyzed and standardized data.

[0839] "Expression" refers to content that is generated using an information scenario and presented to the user through visual and auditory elements.

[0840] "Emotional state" refers to the current psychological and emotional condition of the user, based on their facial expressions and actions.

[0841] "Dynamic adjustment" means changing the content and pace in real time according to the user's emotional state.

[0842] "Feedback" refers to responses and information provided based on the user's experiences and evaluations, and is used as material for improving the system.

[0843] To realize this invention, a server-centric system is built, starting with collecting motion data, analyzing it, and standardizing it. Based on the collected data, the server generates information scenarios and uses these scenarios to form representations. The generated representations are edited to optimize visual and auditory elements and delivered to users via terminals.

[0844] The server utilizes cameras and sensors mounted on the machine to recognize the user's emotional state in real time and analyze the user's facial expressions and movements. Specific examples include the use of emotion analysis libraries such as OpenCV and dlib. This allows for dynamic adjustment of the content of the expressions in response to the user's psychological reactions.

[0845] Users can adjust their work pace and content through the delivered expressions, thereby improving efficiency. For example, if fatigue is detected while a user is performing a certain task, the server can provide expressions intended to promote relaxation, thereby improving the user's work efficiency.

[0846] Furthermore, user feedback is sent to the server via the device. This feedback information will be used to improve expression generation and emotion recognition technologies in the future. Through this iterative learning process, the accuracy of the generative AI model will be improved.

[0847] An example of a prompt message might be, "If the robot shows signs of fatigue during work, generate a refresh video and provide a positive scenario to encourage a break." This allows the server to generate an appropriate scenario based on the instructions, creating a dynamic work environment.

[0848] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0849] Step 1:

[0850] The server collects motion data material through various sensors and cameras. This input data includes videos of work movements and environmental conditions such as temperature. The server converts this data into a digital format and stores it in a database.

[0851] Step 2:

[0852] The server analyzes and standardizes the collected motion data. It uses raw data from the database as input and performs data processing such as noise reduction and data transformation. The standardized data forms the basis for the next processing step.

[0853] Step 3:

[0854] The server generates information scenarios based on standardized data. Utilizing a generative AI model, the algorithm develops the optimal scenario based on the prompt. This output is saved as an operational program and used in the next step.

[0855] Step 4:

[0856] The server generates visual and auditory representations using the generated information scenarios. It performs data calculations to create an effective combination of video and audio, and outputs the results as representation data.

[0857] Step 5:

[0858] The server edits and optimizes the generated representation data. Data processing, such as filtering and adding effects, is performed to balance the visual and auditory elements. This optimized output is then ready to be presented to the user.

[0859] Step 6:

[0860] The terminal delivers optimized content to the user. The user views the presented content and uses it as a guide for their work. Because this delivery is done in real time, a network protocol for delivery is used.

[0861] Step 7:

[0862] The device uses a camera and sensors to recognize the user's emotional state in real time. The data obtained as input is analyzed using libraries such as OpenCV to identify the emotional state. The analysis results are then sent to a server.

[0863] Step 8:

[0864] The server dynamically adjusts the content of its expressions based on sentiment analysis results. Using prompt sentences as a guide, the generative AI model updates the scenario and outputs appropriate content in real time to keep the user interested.

[0865] Step 9:

[0866] Users input subjective experiences and evaluations into their devices through feedback. This input data is sent to a server for improvement and used in future representation generation models.

[0867] Step 10:

[0868] The server uses the collected feedback and emotion recognition results to improve its expression generation model. By analyzing input data and training the generative AI model, it enhances the accuracy of future scenario generation.

[0869] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0870] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0871] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0872] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0873] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0874] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0875] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0876] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0877] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0878] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0879] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0880] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0881] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0882] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0883] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0884] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0885] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0886] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0887] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0888] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0889] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0890] The following is further disclosed regarding the embodiments described above.

[0891] (Claim 1)

[0892] Methods for collecting exercise video material,

[0893] A means for analyzing and standardizing the collected motion video material,

[0894] A means for generating a video scenario based on standardized video material,

[0895] A means of generating video using the generated video scenario,

[0896] A means for editing the generated video and optimizing its visual and auditory elements,

[0897] A means of delivering edited and optimized video to users,

[0898] A means of collecting and analyzing user feedback,

[0899] A means to improve the video generation model based on the analysis results,

[0900] A system that includes this.

[0901] (Claim 2)

[0902] The system according to claim 1, comprising means for combining scenes of failure and scenes of potential success in exercise video material to evoke emotions in the viewer.

[0903] (Claim 3)

[0904] The system according to claim 1, comprising means for personalizing and adjusting the video delivered to the user based on the user's past viewing history and interests.

[0905] "Example 1"

[0906] (Claim 1)

[0907] Means for collecting exercise content,

[0908] A means for analyzing and standardizing the collected motion content,

[0909] A means of generating scenarios based on standardized content,

[0910] A means of generating media using the generated scenario,

[0911] Means for editing the generated media and optimizing its visual and auditory elements,

[0912] A means of delivering content to users using edited and optimized media,

[0913] A means of collecting and analyzing user feedback,

[0914] A means to improve the generative model based on the analysis results,

[0915] A system that includes this.

[0916] (Claim 2)

[0917] The system according to claim 1, comprising means for combining scenes of failure and scenes of potential success in exercise content to evoke emotions in the viewer.

[0918] (Claim 3)

[0919] The system according to claim 1, comprising means for personalizing and adjusting the media delivered to a user based on the user's past usage history and interests.

[0920] "Application Example 1"

[0921] (Claim 1)

[0922] Means for collecting motion media materials,

[0923] A means for analyzing and standardizing the collected kinetic media material,

[0924] A means for generating media scenarios based on standardized data materials,

[0925] A means of generating information using the generated media scenario,

[0926] Means for editing the generated information and optimizing the visual and auditory elements,

[0927] A means of delivering edited and optimized information to users,

[0928] A means of collecting and analyzing user feedback,

[0929] A means to improve the information generation model based on the analysis results,

[0930] A means of generating and distributing personalized exercise recommendation content,

[0931] A system that includes this.

[0932] (Claim 2)

[0933] The system according to claim 1, comprising means for combining scenes of failure and scenes of potential success in motion media material to evoke emotions in the viewer.

[0934] (Claim 3)

[0935] The system according to claim 1, comprising means for personalizing information delivered to users and adjusting it based on the user's past viewing history and interests.

[0936] "Example 2 of combining an emotion engine"

[0937] (Claim 1)

[0938] Means for collecting exercise-related data,

[0939] A means for analyzing and standardizing the collected exercise-related data,

[0940] A means for generating a video scenario based on standardized data,

[0941] A means of generating new video using the generated video scenario,

[0942] A means of editing the generated video and adjusting the visual and auditory elements,

[0943] A means of recognizing user emotions in real time and dynamically adjusting the video,

[0944] A means of providing users with edited and adjusted video,

[0945] A means of collecting and analyzing user responses,

[0946] Means for optimizing the system based on the analyzed response,

[0947] A device that includes this.

[0948] (Claim 2)

[0949] The apparatus according to claim 1, comprising means for analyzing the user's facial expressions and movements through an emotion engine for evoking emotions in the viewer.

[0950] (Claim 3)

[0951] The apparatus according to claim 1, comprising means for personalizing the video provided to the user and adjusting it based on the user's observation history and interests.

[0952] "Application example 2 when combining with an emotional engine"

[0953] (Claim 1)

[0954] Means for collecting exercise data material,

[0955] A means for analyzing and standardizing the collected motion data material,

[0956] A means for generating information scenarios based on standardized data materials,

[0957] A means of generating a representation using the generated information scenario,

[0958] Means for editing the generated representation and optimizing its visual and auditory elements,

[0959] A means of delivering content to users using edited and optimized presentation,

[0960] A means for recognizing the user's emotional state in real time and analyzing their facial expressions and movements,

[0961] A means of dynamically adjusting the content of expression based on the results of emotion analysis,

[0962] A means of collecting and analyzing user feedback,

[0963] Based on the analysis results and emotion recognition results, a means to improve the expression generation model,

[0964] A system that includes this.

[0965] (Claim 2)

[0966] The system according to claim 1, comprising means for combining failure scenes and scenes of potential success in motion data material to evoke emotions in the viewer.

[0967] (Claim 3)

[0968] The system according to claim 1, comprising means for personalizing the content delivered to the user and adjusting it based on the user's past viewing history and interests. [Explanation of Symbols]

[0969] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Methods for collecting exercise video material, A means for analyzing and standardizing the collected motion video material, A means for generating a video scenario based on standardized video material, A means of generating video using the generated video scenario, A means for editing the generated video and optimizing its visual and auditory elements, A means of delivering edited and optimized video to users, A means of collecting and analyzing user feedback, A means to improve the video generation model based on the analysis results, A system that includes this.

2. The system according to claim 1, comprising means for combining scenes of failure and scenes of potential success in exercise video material to evoke emotions in the viewer.

3. The system according to claim 1, comprising means for personalizing the video delivered to the user and adjusting it based on the user's past viewing history and interests.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A