system

The system addresses the need for specific feedback and eco-friendly options in winter sports by analyzing user data to provide personalized technical improvements and sustainable activity suggestions.

JP2026070905APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Winter sports enthusiasts face challenges in obtaining specific feedback for technical improvement and accurate guidance, and there is a lack of easy access to eco-friendly options, hindering sustainable sports participation.

Method used

A system that analyzes user-captured video and audio data to provide personalized technical feedback and environmentally relevant recommendations, using generative AI models to suggest improvements and eco-friendly activities.

Benefits of technology

Enables users to enhance their sports skills while considering environmental impact, offering tailored suggestions and recommendations for sustainable participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070905000001_ABST
    Figure 2026070905000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for identifying areas for improvement in the technology by receiving and analyzing video data captured by the user, A means of analyzing voice data and generating new action suggestions based on user comments, A means of providing recommendation information related to environmental considerations based on the analysis results, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Winter sports enthusiasts have limited means of obtaining specific feedback for technical improvement and at the same time have difficulty receiving accurate guidance in learning new techniques and tricks. Also, as awareness of environmental issues increases, it is difficult to easily know eco-friendly options, and further promotion of sustainable sports participation is required. Therefore, there is a need to provide a system that can solve these problems and simultaneously achieve technical improvement, learning support, and environmental consideration.

Means for Solving the Problems

[0005] This invention includes means for receiving and analyzing video data captured by a user to identify specific areas for technical improvement. It also includes means for analyzing user comments based on audio data to generate new action suggestions. Furthermore, it combines these with means for providing environmentally relevant recommendation information based on the analysis results, aiming to support technical improvement, new challenges, and sustainable participation in sports. This system allows users to balance practice aimed at technical improvement with environmental awareness.

[0006] A "user" is someone who uses the system to receive technological improvements and new suggestions.

[0007] "Video data" refers to video files that record the user's movements while skiing or snowboarding.

[0008] "Analysis" is the process of processing collected video and audio data to extract meaningful information.

[0009] "Technical improvements" refer to specific suggestions for enhancing the user's sports performance.

[0010] "Audio data" refers to audio files in which users record their own comments and feedback.

[0011] A "new behavioral suggestion" is a suggestion of a new technology or trick that users can try out.

[0012] "Recommendations related to environmental considerations" refers to information about eco-friendly ski resorts and environmentally conscious options.

[0013] "Receiving" refers to the server retrieving data sent by the user. [Brief explanation of the drawing]

[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is a system designed to simultaneously improve technical skills and environmental awareness in winter sports. The system allows users to upload video and audio data captured using their individual terminals to a server, where the data is analyzed to provide personalized feedback and suggestions.

[0036] The server takes the lead in receiving video data and uses a multimodal generative AI model to analyze the user's technical movements. This analysis extracts specific performance data, such as the position of the center of gravity and the consistency of movement. Based on this, improvements and new technology proposals are made.

[0037] In addition, the server converts the audio data into text and generates specific suggestions, including new tricks and technical challenges, based on the user's comments. These suggestions are customized to the user's skill level and interests.

[0038] Furthermore, the server accesses a database related to environmental considerations and provides information recommending eco-friendly ski resorts and activities suitable for the user. This allows users to enjoy sports in a sustainable way, not just improve their skills.

[0039] The device presents these analysis results and suggestions through a user interface, displaying feedback in a visually easy-to-understand format. Based on the feedback, users can create practice plans aimed at improving their next performance or consider environmentally conscious options.

[0040] For example, if a user uploads a video of themselves snowboarding along with a voice comment saying, "I want to improve the stability of my jumps," the server will analyze data on the user's posture and speed during the jumps from the video and suggest improvements such as shifting their center of gravity forward. It will also recommend ski resorts that offer environmentally friendly facilities, providing options for future visits. The device will display this information and feedback to the user in an easy-to-understand way, allowing them to use it to improve their next activity.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] Users can use their smartphones or cameras to film videos of themselves skiing or snowboarding, and record voice comments about their performance as needed.

[0044] Step 2:

[0045] The user launches a dedicated application and selects the captured video and audio data. The user then presses the upload button to send this data to the server. The data may be compressed during transmission depending on network conditions.

[0046] Step 3:

[0047] The server receives data from the user. After receiving the data, the server performs preprocessing to convert the video and audio data into a parseable format. This includes processing each frame of the video and denoising the audio data.

[0048] Step 4:

[0049] The server analyzes the video data using a multimodal generative AI model. The AI ​​analyzes the motion frames, extracts the user's technical elements (e.g., center of gravity, turn angle, speed changes), and identifies areas for technical improvement based on this.

[0050] Step 5:

[0051] The server converts the audio data into text and analyzes the user's comments. This generates suggestions for new tricks and actions tailored to the user's interests and goals.

[0052] Step 6:

[0053] Based on the analysis results, the server recommends suitable ski resorts and activities to the user from a database of eco-friendly information. This information supports sustainable participation in sports.

[0054] Step 7:

[0055] The device receives feedback information from the server. The feedback is displayed in a visual and interactive format that is easy for the user to understand.

[0056] Step 8:

[0057] Users review feedback information to improve technology and plan new activities. They also consider the eco-friendly options offered and use that information to inform their next visit.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] In modern sports activities, there are limited systems that simultaneously support individual skill improvement and increased environmental awareness. In particular, in winter sports, there is a need for users to objectively evaluate their own skills, obtain concrete guidance for improvement, and acquire appropriate information to enjoy the activity in a sustainable manner. This invention aims to meet these needs by developing a system that provides feedback on users' technical actions and recommends environmentally conscious activities.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes means for receiving and analyzing video data captured by the user to identify areas for improvement in exercise technique, means for analyzing audio data to generate new action suggestions based on the user's intentions and comments, and means for providing recommendation information related to environmental protection based on the analysis results. As a result, the user can receive an objective evaluation of their technical performance and obtain specific feedback that leads to improvement, as well as receive specific information to help them choose environmentally conscious activities.

[0063] A "user" refers to a person who uses this system to improve their skills and engage in sports activities while being mindful of the environment.

[0064] "Video data" refers to video footage of sports activities filmed by users, and the analysis of this data enables technological improvements.

[0065] "Voice data" refers to audio information, including comments and instructions generated by the user, and new action suggestions are generated based on its content.

[0066] A "server" refers to a core processing unit that receives video and audio data, provides computing resources for analysis, and returns feedback to the user.

[0067] "Analysis" refers to the process of processing received data to extract useful information and generating suggestions or feedback based on that information.

[0068] A "generative AI model" refers to artificial intelligence technology that supports analytical processing, and is responsible for evaluating user behavior and generating appropriate feedback.

[0069] A "prompt sentence" refers to a sentence that is input into the AI ​​model based on the transcribed information of the audio data, and it forms the basis for feedback generation.

[0070] "Environmental protection-related recommendations" refer to information about environmentally friendly locations and activities that are useful for users to enjoy sports in a sustainable way.

[0071] "User interface" refers to the method by which users receive analysis results and suggestions from a server and view the information visualized in an easy-to-understand format.

[0072] This invention is a system that allows users to improve their winter sports activities while also making environmentally conscious choices. Specifically, users capture video and audio data using a smartphone or other camera-equipped device. This data is then uploaded to a server via an application.

[0073] The server utilizes a generative AI model to analyze the received video data. This model evaluates, for example, changes in the user's center of gravity and the accuracy of jumps frame by frame. Machine learning libraries such as PyTorch and TENSORFLOW® are used for the analysis, enabling precise data analysis.

[0074] For audio data, speech recognition technologies such as Google® Cloud Speech-to-Text API are used to convert it to text. The converted text is then input into a generative AI model, which generates specific feedback and new suggestions based on the user's intent.

[0075] A terminal equipped with a user interface receives analysis results and suggestions from the server and displays them in a format that is easy for the user to understand. For example, it presents suggestions for improvement derived from video data and recommendations for environmentally friendly ski resorts in an organized dashboard format.

[0076] For example, if a user uploads a video of themselves snowboarding and adds a voice comment saying, "I want to improve the stability of my jumps," the system will analyze the user's posture and speed during the jump. As a result, it will generate suggestions such as, "It would be good to move your center of gravity a little further forward when you jump."

[0077] An example of a prompt message would be: "Based on the following video and audio commentary, generate technical feedback on winter sports. Please include specific improvement suggestions, taking into account the video analysis results."

[0078] By using this system, users can receive technical feedback while also making environmentally conscious choices, leading to a more fulfilling sports experience.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] Users use smartphones or camera-equipped devices to record videos and audio commentary of winter sports. The input consists of video and audio data, which will serve as foundational data for later analysis. Specifically, users press a button to record and, if necessary, leave voice commentary.

[0082] Step 2:

[0083] The device uploads video and audio data captured by the user to a server via a dedicated app. The input consists of video and audio data from the user, and these are sent to the server as output. Specifically, the upload begins when the user taps the send button on the app screen.

[0084] Step 3:

[0085] The server receives the uploaded video data and activates the analysis module. The input is video data sent from the terminal, and the output is numerical data related to the user's technical performance. Using a generative AI model, the analysis is performed frame by frame to evaluate the user's center of gravity and consistency of movement. Specifically, the server decomposes the video into frames and evaluates each frame sequentially.

[0086] Step 4:

[0087] The server starts a speech recognition engine to convert audio data into text. The input is audio data, and the output is the user's comments in text format. Specifically, the server passes the audio data through the speech recognition engine, then converts it to text format to prepare it for analysis.

[0088] Step 5:

[0089] The server integrates the results of video analysis and audio-to-text conversion, and generates feedback using a generative AI model. The input is the analyzed video and text data, and the output is specific technical improvements and new suggestions. Based on this, suggestions such as "move the center of gravity forward when jumping" are created. Specifically, both sets of data are analyzed in an integrated manner, and prompt sentences are input into the generative AI model to obtain feedback.

[0090] Step 6:

[0091] The server provides users with environmentally friendly recommendations. Input consists of analyzed video and environmental databases, while output is information on eco-friendly activities suitable for the user. Specifically, the server retrieves information from relevant databases based on the user's location and activity data, and then makes suggestions to the user.

[0092] Step 7:

[0093] The device receives feedback and recommendation information sent from the server and displays it visually. Input consists of feedback data and recommendation information from the server, while output is displayed in a user-friendly format. Specifically, the device displays a dashboard screen, visualizing analysis results and suggestions as icons and graphs. Based on this, the user can plan their next practice session.

[0094] (Application Example 1)

[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] Winter sports enthusiasts often have difficulty obtaining appropriate feedback for improving their technique, and they also lack sufficient information regarding post-sports nutrition. Furthermore, they tend to be indifferent to information concerning environmental considerations. Therefore, there is a need to provide a system that allows users to enjoy sports in a way that promotes both technical improvement and environmental considerations.

[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0098] In this invention, the server includes means for receiving and analyzing video and image data captured by the user to identify areas for improvement in technique, means for analyzing audio data and generating new action suggestions based on the user's comments, means for providing recommendation information related to environmental considerations based on the analysis results, and means for recommending highly nutritious menus based on exercise data. This makes it possible for users to improve their sports technique while also selecting appropriate nutritional support and environmentally conscious activities after sports.

[0099] "Users" refer to individuals who participate in winter sports or those who aim to improve their skills in those sports.

[0100] "Video and image data" refers to video and image files taken by users during winter sports activities.

[0101] "Analysis" is the act of analyzing specific data and extracting useful information from it.

[0102] "Technical improvements" refer to specific changes or points of instruction needed to improve current performance.

[0103] "Audio data" refers to audio data such as comments and explanations recorded by the user.

[0104] "Action suggestions" are specific guidelines or advice that indicate what action a user should take next.

[0105] "Environmentally conscious recommendation information" refers to information that suggests eco-friendly actions and places, encouraging choices that take sustainability into consideration.

[0106] "Exercise data" refers to data that includes activity records generated when a user participates in winter sports.

[0107] "Recommending nutritious menus" means suggesting meal options that contain the necessary nutrients based on the user's activity level and physical condition.

[0108] This invention is a system that enables winter sports participants to achieve both skill improvement and environmental consideration. Users collect video and audio data during exercise using a mobile device. This data is uploaded to a server. The server analyzes this data using a generative AI model.

[0109] The server analyzes video data and extracts specific performance metrics related to the user's skills. For example, it can obtain data such as the position of the center of gravity and the consistency of movement. Audio data is converted into text based on the user's preferences and comments, and processed through a generative AI model that suggests new actions and technical challenges.

[0110] Furthermore, the server analyzes the user's activity level based on their exercise data and suggests nutritional supplements needed after exercise. This includes a process that recommends nutritious menus using local ingredients. Eco-friendly activities and locations are also suggested, enabling users to enjoy sports in a sustainable way.

[0111] As a concrete example, after a user goes skiing, the application asks, "What kind of training should I do the next day?" The system then provides new training methods and nutritional suggestions based on the user's previous activity data. An example of a prompt message would be, "Generate training suggestions and a meal plan based on the user's recent exercise data."

[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0113] Step 1:

[0114] Users collect video and audio data of their winter sports activities using their mobile devices. This involves recording video using the device's camera and recording audio commentary using the microphone. This data is stored on the device and prepared for uploading to the server.

[0115] Step 2:

[0116] The terminal uploads collected video and audio data to the server. Data transfer takes place via an internet connection, and the server receives this data. The input data is originally in various formats, but it is converted to a unified format here.

[0117] Step 3:

[0118] The server analyzes the received video data. Using a generative AI model, it extracts metrics related to user performance. In this process, it analyzes data such as the center of gravity and motion patterns to identify areas for improvement. The input is video data, and the output is the analyzed performance data.

[0119] Step 4:

[0120] The server converts audio data to text and generates new action suggestions from user comments. It uses speech recognition software to convert speech to text and analyzes its content to understand the user's intent and requests. The input is audio data, and the output is action suggestions in text format.

[0121] Step 5:

[0122] The server provides environmentally conscious recommendations and nutritious menu suggestions based on analyzed performance data and action suggestions. This includes eco-friendly options and meal plans tailored to the user's activity level. Inputs are analyzed data and action suggestions, and output is provided as recommendations.

[0123] Step 6:

[0124] The terminal displays feedback received from the server to the user. Through the user interface, analysis results and suggestions are presented in an easy-to-understand format. Based on this, the user decides on their next exercise plan and meal choices. The input is feedback from the server, and the output is the notification content to the user.

[0125] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0126] This invention is a system that combines a new emotion engine to improve skills and environmental awareness in winter sports. Users upload video and audio data of themselves skiing or snowboarding to a server using their own devices. By analyzing this data, the system evaluates the user's performance in detail and provides specific feedback for skill improvement.

[0127] The server receives video data in real time and analyzes the user's movements using an AI model. This allows it to evaluate detailed technical elements such as turn angles and balance, and identify specific areas for improvement. In addition, it extracts the user's self-assessment from audio data and generates personalized action suggestions.

[0128] The newly integrated emotion engine recognizes emotions from the user's voice and facial expressions, contributing to improved analysis accuracy. Specifically, the emotion engine determines whether the user is enjoying themselves or experiencing difficulties, and adjusts the feedback and suggestions provided to match the user's emotional state. This emotional information is also applied to suggest actions that enhance emotional satisfaction.

[0129] Furthermore, recommendations related to environmental considerations are provided in a personalized manner that takes into account the user's emotional state. For example, a user seeking relaxation might be recommended a ski resort that offers eco-friendly services in a quiet environment, using emotional information to provide more individualized suggestions.

[0130] For example, if the emotion engine recognizes a user's voice comment saying, "I'm struggling with my snowboard turns," along with a confused expression, the server will provide detailed suggestions for improving their turning technique and even suggest simple tricks to help them maintain balance. In addition, information on eco-friendly ski resorts where the user can practice while having fun will also be provided.

[0131] Thus, the present invention aims to further enhance the winter sports experience by combining flexible feedback that responds to the user's emotions with sustainable options.

[0132] The following describes the processing flow.

[0133] Step 1:

[0134] Users can use their smartphones or cameras to film videos and videos of themselves skiing or snowboarding, and record their thoughts and challenges in audio as needed.

[0135] Step 2:

[0136] Users select video and audio data through a dedicated application and upload them to the server. The upload is performed using a secure and efficient data transfer protocol.

[0137] Step 3:

[0138] The server receives the uploaded data. Upon receipt, the server verifies the data format and performs any necessary preprocessing. This includes frame extraction and audio cleaning.

[0139] Step 4:

[0140] The server supplies video data to a multimodal AI model for technical motion analysis. Here, technical metrics such as turn precision, balance, and jump height are extracted.

[0141] Step 5:

[0142] The server analyzes the voice data to identify the user's intentions and challenges. This allows for the generation of specific technical improvement suggestions based on the user's interests.

[0143] Step 6:

[0144] The server uses an emotion engine to analyze the user's emotions from video and audio data. It uses facial expressions and tone of voice to identify the user's current emotional state (e.g., satisfaction, impatience, excitement).

[0145] Step 7:

[0146] The server customizes feedback and action suggestions based on the analysis results and the user's emotional state. The tone and level of detail of the suggestions are adjusted according to the emotional state.

[0147] Step 8:

[0148] The server generates personalized recommendations related to eco-friendly activities based on emotional information. For example, it might recommend ski resorts with a relaxing and quiet environment.

[0149] Step 9:

[0150] The device receives feedback and suggestions from the server and presents them to the user in a visual and interactive format. This allows the user to intuitively understand the feedback and plan their next steps.

[0151] (Example 2)

[0152] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0153] In winter sports, it is difficult to provide a highly satisfying experience because there is a lack of means to provide environmentally conscious recommendation information that is tailored to the emotional state of the user, in addition to improving their skills.

[0154] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0155] In this invention, the server includes means for receiving video information, analyzing technical operations using a generation AI model to identify areas for improvement, analyzing audio information, recognizing the user's emotional state to generate new action suggestions, providing environmentally conscious recommendation information that takes the emotional state into account, and recognizing the user's emotions using an emotion engine to personalize feedback. This makes it possible to provide users with personalized technical improvement feedback and recommendation information that takes sustainability into consideration.

[0156] "Video and image information" refers to data including images and actions captured by the user during winter sports activities.

[0157] A "generative AI model" refers to an artificial intelligence algorithm used to analyze received video and image information and identify areas for improvement to enhance specific technologies.

[0158] "Audio information" refers to data that includes verbal comments and impressions made by users during or after participating in winter sports.

[0159] An "emotion engine" refers to a technological element used to recognize and analyze a user's emotions from their voice and facial expressions, and to supplement the analysis results.

[0160] "Environmentally conscious recommendation information" refers to information about facilities and services that allow users to enjoy winter sports in a more sustainable way, while taking into account their emotional state.

[0161] This invention is a system for improving winter sports techniques and providing environmentally conscious recommendation information tailored to user emotions. Users use smartphones or action cameras as terminals to collect video and audio information while skiing or snowboarding. Users upload this data to a server via their terminals.

[0162] The server inputs video data into an AI model to analyze the user's technical movements. This model evaluates the angle and speed of turns, as well as balance, and identifies specific areas for improvement. The server also analyzes audio data using natural language processing technology to extract the user's self-assessment and problems. Furthermore, it uses an emotion engine to analyze the user's facial expressions and tone of voice to recognize the user's emotional state.

[0163] Based on the analysis results, the server generates feedback for technical improvements tailored to the user, and also provides environmentally conscious recommendations that match their emotional state. For example, a user who wants to relax might be recommended an eco-friendly ski resort with a quiet environment. The feedback and recommendations are sent to the user's device, where they can review them and use them to plan their next activities.

[0164] For example, if a user provides a voice comment such as, "I'm having trouble with my snowboard turns," and the emotion engine recognizes from their facial expression that they are confused, the server will provide specific advice on improving their turning technique and also suggest eco-friendly ski resorts that the user can enjoy.

[0165] Examples of prompt statements include the following:

[0166] "Please give me some specific advice on how to improve my snowboarding turning technique."

[0167] "Please recommend an eco-friendly ski resort."

[0168] This invention allows users to improve their winter sports skills while gaining an emotionally conscious and sustainable experience.

[0169] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0170] Step 1:

[0171] Users collect video and audio information

[0172] The user uses a device to record video and audio information during winter sports activities, along with comments and emotional expressions. This collects data necessary for evaluating technical actions and emotional states. The input for this step is the user's actions and audio recordings in the field, and the output is the corresponding digital data.

[0173] Step 2:

[0174] User uploads data to the server

[0175] The user sends video and audio information collected via their device to the server. The user uploads the data using a dedicated app or web interface. The input in this step is digital data stored on the user's device, and the output is the information received by the server.

[0176] Step 3:

[0177] The server analyzes the video information.

[0178] The server inputs the received video information into a generating AI model, which analyzes technical actions such as turn angle, speed, and balance. The AI ​​model uses advanced algorithms to detect detailed technical elements and evaluate the user's performance quality. The input for this step is video information, and the output is the technical evaluation result.

[0179] Step 4:

[0180] The server analyzes the audio information.

[0181] The server analyzes the audio information using natural language processing technology. It extracts self-assessments and problems from user comments, using this as foundational data to generate action suggestions. The input for this step is audio information, and the output is the results of the self-assessment and problem extraction.

[0182] Step 5:

[0183] The server recognizes emotions.

[0184] The server uses an emotion engine to recognize the user's emotional state from video and audio tones. This analysis determines the emotional situation, such as whether the user is enjoying themselves or is confused. The input for this step is video and audio tones, and the output is the result of the emotion judgment.

[0185] Step 6:

[0186] The server generates feedback and recommendation information.

[0187] The server generates user-appropriate feedback based on technical evaluations, self-assessment results, and emotional judgments. Specific advice for technical improvement and emotionally responsive, environmentally conscious recommendations are provided. The input for this step is the entire analysis result, and the output is user feedback and recommendations.

[0188] Step 7:

[0189] Users receive feedback

[0190] The user receives and confirms feedback and recommendations sent from the server through the terminal interface. Based on the received information, the user plans their next activity. The input for this step is data from the server, and the output is the information confirmed by the user.

[0191] (Application Example 2)

[0192] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0193] In winter sports, there are problems such as a lack of detailed feedback for improving technique, a lack of personalized feedback tailored to the user's emotions, and limited information on service selection that takes environmental considerations into account. This project aims to solve these technical challenges.

[0194] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0195] This invention includes a server that includes means for receiving and analyzing video data captured by the user to identify areas for technological improvement, means for analyzing audio data and generating new action suggestions based on the user's comments, an emotion engine that recognizes the user's emotions and provides personalized feedback based on the analysis results, and means for providing recommendation information related to environmental considerations based on the analysis results and recognized emotions. This enables users enjoying winter sports to receive detailed technological feedback, personalized suggestions tailored to their emotions, and information on eco-friendly options.

[0196] "Communication equipment" is a general term for information terminals used by users to transmit data to external data processing devices.

[0197] A "data processing device" refers to a computer system that analyzes video and audio data received from users and generates action suggestions and feedback.

[0198] "Video data" refers to video files that record user activity and is used for analysis to identify areas for technical improvement.

[0199] "Voice data" refers to audio files that record the user's voice and contain information useful for emotion recognition and generating action suggestions.

[0200] An "emotion engine" is a software module that recognizes a user's emotional state from voice and visual data and personalizes suggestions and feedback based on that state.

[0201] "Eco-friendly options" refer to information about sustainable services and products that take environmental considerations into account, and are included in the recommendation information provided to users.

[0202] "Feedback" is a process of providing information that indicates areas for improvement and recommended actions based on user activity, and is particularly aimed at improving users' technical skills and the quality of their experience.

[0203] This invention provides a system for offering winter sports enthusiasts an experience that balances technological advancement with environmental consideration. The system is started when a user captures video data using a smartphone and uploads it to a data processing device via a communication device.

[0204] The server analyzes the received video data using computer vision tools (e.g., OpenCV) to extract the technical characteristics of the user's movements. This information includes technical elements such as the angle of turns and balance. Furthermore, the audio data is analyzed using speech recognition tools (e.g., Google Speech-to-Text, emotion recognition API) to identify the user's comments and emotional state.

[0205] The emotion engine recognizes the user's emotions from analyzed voice and facial expression data and generates personalized feedback and action suggestions based on that information. For example, if a user comments that they are "struggling with turns," the server will provide specific action steps to improve their turning technique and also offer practice suggestions to improve their balance. Furthermore, depending on the user's emotional state, information about eco-friendly facilities will be recommended for users seeking relaxation.

[0206] An example prompt would be, "Analyze the user's skiing video and audio data, and provide personalized information based on technical evaluation and emotion recognition." This system aims to provide users with a fulfilling sports experience by combining detailed technical checks in winter sports with emotionally tailored support, offering sustainable options.

[0207] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0208] Step 1:

[0209] The system uses a device to record video and audio data captured by the user during winter sports activities. Input is the video and audio captured by the user, and output is a data file stored on the device. Specifically, the device's camera and microphone are used to record footage of activities such as riding ski lifts or skiing.

[0210] Step 2:

[0211] The terminal uploads recorded video and audio data to the server via a communication device. The input is the video and audio data files stored on the terminal, and the output is the data transferred to the server. Specifically, the terminal transmits data wirelessly via a dedicated application.

[0212] Step 3:

[0213] The server analyzes the received video data using computer vision tools (e.g., OpenCV). The input is the video data uploaded to the server, and the output is the technical features obtained through the analysis (e.g., turn angle, speed, balance). Specifically, the system processes the video data frames sequentially and calculates various metrics of the user's movement.

[0214] Step 4:

[0215] The server receives audio data, converts it to text using a speech recognition tool (e.g., Google Speech-to-Text), and then identifies the emotional state using an emotion recognition API. The input is audio data uploaded to the server, and the output is the user's comments and the recognized emotion information. Specifically, the process converts the audio data to text and extracts emotions from the content and intonation.

[0216] Step 5:

[0217] The server integrates analyzed technical features and emotional information, and uses a generative AI model to generate user-appropriate feedback and action suggestions. The input is technical features and emotional information, and the output is personalized feedback and action suggestions. Specifically, the prompt message "Analyze the user's skiing video and audio data, and provide personalized information based on technical evaluation and emotional recognition" is input to the AI ​​model, and the generated suggestions are created.

[0218] Step 6:

[0219] The server generates feedback and action suggestions, which are then sent back to the terminal and presented to the user. The input is the generated suggestions, and the output is the feedback message that the user views on the terminal. Specifically, the information is displayed to the user as text messages or video clips through a dedicated application.

[0220] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0221] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0222] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0223] [Second Embodiment]

[0224] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0225] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0226] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0227] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0228] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0229] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0230] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0231] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0232] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0233] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0234] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0235] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0236] This invention is a system designed to simultaneously improve technical skills and environmental awareness in winter sports. The system allows users to upload video and audio data captured using their individual terminals to a server, where the data is analyzed to provide personalized feedback and suggestions.

[0237] The server takes the lead in receiving video data and uses a multimodal generative AI model to analyze the user's technical movements. This analysis extracts specific performance data, such as the position of the center of gravity and the consistency of movement. Based on this, improvements and new technology proposals are made.

[0238] In addition, the server converts the audio data into text and generates specific suggestions, including new tricks and technical challenges, based on the user's comments. These suggestions are customized to the user's skill level and interests.

[0239] Furthermore, the server accesses a database related to environmental considerations and provides information recommending eco-friendly ski resorts and activities suitable for the user. This allows users to enjoy sports in a sustainable way, not just improve their skills.

[0240] The device presents these analysis results and suggestions through a user interface, displaying feedback in a visually easy-to-understand format. Based on the feedback, users can create practice plans aimed at improving their next performance or consider environmentally conscious options.

[0241] For example, if a user uploads a video of themselves snowboarding along with a voice comment saying, "I want to improve the stability of my jumps," the server will analyze data on the user's posture and speed during the jumps from the video and suggest improvements such as shifting their center of gravity forward. It will also recommend ski resorts that offer environmentally friendly facilities, providing options for future visits. The device will display this information and feedback to the user in an easy-to-understand way, allowing them to use it to improve their next activity.

[0242] The following describes the processing flow.

[0243] Step 1:

[0244] Users can use their smartphones or cameras to film videos of themselves skiing or snowboarding, and record voice comments about their performance as needed.

[0245] Step 2:

[0246] The user launches a dedicated application and selects the captured video and audio data. The user then presses the upload button to send this data to the server. The data may be compressed during transmission depending on network conditions.

[0247] Step 3:

[0248] The server receives data from the user. After receiving the data, the server performs preprocessing to convert the video and audio data into a parseable format. This includes processing each frame of the video and denoising the audio data.

[0249] Step 4:

[0250] The server analyzes the video data using a multimodal generative AI model. The AI ​​analyzes the motion frames, extracts the user's technical elements (e.g., center of gravity, turn angle, speed changes), and identifies areas for technical improvement based on this.

[0251] Step 5:

[0252] The server converts the audio data into text and analyzes the user's comments. This generates suggestions for new tricks and actions tailored to the user's interests and goals.

[0253] Step 6:

[0254] Based on the analysis results, the server recommends suitable ski resorts and activities to the user from a database of eco-friendly information. This information supports sustainable participation in sports.

[0255] Step 7:

[0256] The device receives feedback information from the server. The feedback is displayed in a visual and interactive format that is easy for the user to understand.

[0257] Step 8:

[0258] Users review feedback information to improve technology and plan new activities. They also consider the eco-friendly options offered and use that information to inform their next visit.

[0259] (Example 1)

[0260] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0261] In modern sports activities, there are limited systems that simultaneously support individual skill improvement and increased environmental awareness. In particular, in winter sports, there is a need for users to objectively evaluate their own skills, obtain concrete guidance for improvement, and acquire appropriate information to enjoy the activity in a sustainable manner. This invention aims to meet these needs by developing a system that provides feedback on users' technical actions and recommends environmentally conscious activities.

[0262] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0263] In this invention, the server includes means for receiving and analyzing video data captured by the user to identify areas for improvement in exercise technique, means for analyzing audio data to generate new action suggestions based on the user's intentions and comments, and means for providing recommendation information related to environmental protection based on the analysis results. As a result, the user can receive an objective evaluation of their technical performance and obtain specific feedback that leads to improvement, as well as receive specific information to help them choose environmentally conscious activities.

[0264] A "user" refers to a person who uses this system to improve their skills and engage in sports activities while being mindful of the environment.

[0265] "Video data" refers to video footage of sports activities filmed by users, and the analysis of this data enables technological improvements.

[0266] "Voice data" refers to audio information, including comments and instructions generated by the user, and new action suggestions are generated based on its content.

[0267] A "server" refers to a core processing unit that receives video and audio data, provides computing resources for analysis, and returns feedback to the user.

[0268] "Analysis" refers to the process of processing received data to extract useful information and generating suggestions or feedback based on that information.

[0269] A "generative AI model" refers to artificial intelligence technology that supports analytical processing, and is responsible for evaluating user behavior and generating appropriate feedback.

[0270] A "prompt sentence" refers to a sentence that is input into the AI ​​model based on the transcribed information of the audio data, and it forms the basis for feedback generation.

[0271] "Environmental protection-related recommendations" refer to information about environmentally friendly locations and activities that are useful for users to enjoy sports in a sustainable way.

[0272] "User interface" refers to the method by which users receive analysis results and suggestions from a server and view the information visualized in an easy-to-understand format.

[0273] This invention is a system that allows users to improve their winter sports activities while also making environmentally conscious choices. Specifically, users capture video and audio data using a smartphone or other camera-equipped device. This data is then uploaded to a server via an application.

[0274] The server utilizes a generative AI model to analyze the received video data. This model evaluates, for example, changes in the user's center of gravity and the accuracy of jumps frame by frame. Machine learning libraries such as PyTorch and TensorFlow are used for the analysis, enabling precise data analysis.

[0275] For audio data, speech recognition technologies such as the Google Cloud Speech-to-Text API are used to convert it to text. The converted text is then input into a generative AI model, which generates specific feedback and new suggestions based on the user's intent.

[0276] A terminal equipped with a user interface receives analysis results and suggestions from the server and displays them in a format that is easy for the user to understand. For example, it presents suggestions for improvement derived from video data and recommendations for environmentally friendly ski resorts in an organized dashboard format.

[0277] For example, if a user uploads a video of themselves snowboarding and adds a voice comment saying, "I want to improve the stability of my jumps," the system will analyze the user's posture and speed during the jump. As a result, it will generate suggestions such as, "It would be good to move your center of gravity a little further forward when you jump."

[0278] An example of a prompt message would be: "Based on the following video and audio commentary, generate technical feedback on winter sports. Please include specific improvement suggestions, taking into account the video analysis results."

[0279] By using this system, users can receive technical feedback while also making environmentally conscious choices, leading to a more fulfilling sports experience.

[0280] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0281] Step 1:

[0282] The user uses a smartphone or a device with a camera to shoot videos and voice comments of winter sports. The input obtained is video image data and audio data. These are the basic data to be analyzed later. As a specific operation, press the shooting button to record, and if necessary, speak to leave comments.

[0283] Step 2:

[0284] The terminal uploads the video image data and audio data shot by the user to the server via a dedicated app. The input obtained is the video image data and audio data from the user, and the output is that they are sent to the server. Specifically, the upload is started by the user tapping the send button on the app screen.

[0285] Step 3:

[0286] The server receives the uploaded video image data and activates the analysis module. The input is the video image data sent from the terminal, and the output is numerical data regarding the user's technical operation performance. Using the generated AI model, analysis is performed for each frame to evaluate the user's center of gravity position and the consistency of movements. As a specific operation, the server decomposes the video into frames and evaluates each frame in order.

[0287] Step 4:

[0288] The server activates the speech recognition engine to convert the audio data into text. The input is the audio data, and the output is the texturized user comments. As a specific operation, the server passes the audio data through the speech recognition engine and then converts it into text format to prepare for analysis.

[0289] Step 5:

[0290] The server integrates the results of video analysis and audio-to-text conversion, and generates feedback using a generative AI model. The input is the analyzed video and text data, and the output is specific technical improvements and new suggestions. Based on this, suggestions such as "move the center of gravity forward when jumping" are created. Specifically, both sets of data are analyzed in an integrated manner, and prompt sentences are input into the generative AI model to obtain feedback.

[0291] Step 6:

[0292] The server provides users with environmentally friendly recommendations. Input consists of analyzed video and environmental databases, while output is information on eco-friendly activities suitable for the user. Specifically, the server retrieves information from relevant databases based on the user's location and activity data, and then makes suggestions to the user.

[0293] Step 7:

[0294] The device receives feedback and recommendation information sent from the server and displays it visually. Input consists of feedback data and recommendation information from the server, while output is displayed in a user-friendly format. Specifically, the device displays a dashboard screen, visualizing analysis results and suggestions as icons and graphs. Based on this, the user can plan their next practice session.

[0295] (Application Example 1)

[0296] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0297] Winter sports enthusiasts often have difficulty obtaining appropriate feedback for improving their technique, and they also lack sufficient information regarding post-sports nutrition. Furthermore, they tend to be indifferent to information concerning environmental considerations. Therefore, there is a need to provide a system that allows users to enjoy sports in a way that promotes both technical improvement and environmental considerations.

[0298] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0299] In this invention, the server includes means for receiving and analyzing video and image data captured by the user to identify areas for improvement in technique, means for analyzing audio data and generating new action suggestions based on the user's comments, means for providing recommendation information related to environmental considerations based on the analysis results, and means for recommending highly nutritious menus based on exercise data. This makes it possible for users to improve their sports technique while also selecting appropriate nutritional support and environmentally conscious activities after sports.

[0300] "Users" refer to individuals who participate in winter sports or those who aim to improve their skills in those sports.

[0301] "Video and image data" refers to video and image files taken by users during winter sports activities.

[0302] "Analysis" is the act of analyzing specific data and extracting useful information from it.

[0303] "Technical improvements" refer to specific changes or points of instruction needed to improve current performance.

[0304] "Audio data" refers to audio data such as comments and explanations recorded by the user.

[0305] An "action proposal" is specific guidance or advice indicating what actions a user should take next.

[0306] "Recommendation information related to environmental consideration" is information that proposes eco-friendly actions and locations, and encourages choices considering sustainability.

[0307] "Exercise data" is data that includes activity records generated when a user engages in winter sports.

[0308] "Recommend high-nutrition menus" means proposing dietary choices that contain the necessary nutrients based on the user's exercise volume and physical condition.

[0309] This invention is a system for winter sports participants to achieve both technical improvement and environmental consideration. The user uses a mobile device to collect moving image data and audio data during exercise. This data is uploaded to a server. The server analyzes this data using a generative AI model.

[0310] The server analyzes the moving image data and extracts specific performance indicators related to the user's skills. For example, data such as the position of the center of gravity and the consistency of movements can be obtained. The audio data is text-converted based on the user's preferences and comments, and processed through a generative AI model that proposes new actions and technical challenges.

[0311] Furthermore, the server analyzes the activity level obtained from the user's exercise data and proposes nutritional replenishment required after exercise. This includes a process of recommending high-nutrition menus using local ingredients. Eco-friendly activities and locations are also proposed, enabling the user to enjoy sports in a sustainable manner.

[0312] As a concrete example, after a user goes skiing, the application asks, "What kind of training should I do the next day?" The system then provides new training methods and nutritional suggestions based on the user's previous activity data. An example of a prompt message would be, "Generate training suggestions and a meal plan based on the user's recent exercise data."

[0313] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0314] Step 1:

[0315] Users collect video and audio data of their winter sports activities using their mobile devices. This involves recording video using the device's camera and recording audio commentary using the microphone. This data is stored on the device and prepared for uploading to the server.

[0316] Step 2:

[0317] The terminal uploads collected video and audio data to the server. Data transfer takes place via an internet connection, and the server receives this data. The input data is originally in various formats, but it is converted to a unified format here.

[0318] Step 3:

[0319] The server analyzes the received video data. Using a generative AI model, it extracts metrics related to user performance. In this process, it analyzes data such as the center of gravity and motion patterns to identify areas for improvement. The input is video data, and the output is the analyzed performance data.

[0320] Step 4:

[0321] The server converts audio data to text and generates new action suggestions from user comments. It uses speech recognition software to convert speech to text and analyzes its content to understand the user's intent and requests. The input is audio data, and the output is action suggestions in text format.

[0322] Step 5:

[0323] The server provides environmentally conscious recommendations and nutritious menu suggestions based on analyzed performance data and action suggestions. This includes eco-friendly options and meal plans tailored to the user's activity level. Inputs are analyzed data and action suggestions, and output is provided as recommendations.

[0324] Step 6:

[0325] The terminal displays feedback received from the server to the user. Through the user interface, analysis results and suggestions are presented in an easy-to-understand format. Based on this, the user decides on their next exercise plan and meal choices. The input is feedback from the server, and the output is the notification content to the user.

[0326] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0327] This invention is a system that combines a new emotion engine to improve skills and environmental awareness in winter sports. Users upload video and audio data of themselves skiing or snowboarding to a server using their own devices. By analyzing this data, the system evaluates the user's performance in detail and provides specific feedback for skill improvement.

[0328] The server receives video data in real time and analyzes the user's movements using an AI model. This allows it to evaluate detailed technical elements such as turn angles and balance, and identify specific areas for improvement. In addition, it extracts the user's self-assessment from audio data and generates personalized action suggestions.

[0329] The newly integrated emotion engine recognizes emotions from the user's voice and facial expressions, contributing to improved analysis accuracy. Specifically, the emotion engine determines whether the user is enjoying themselves or experiencing difficulties, and adjusts the feedback and suggestions provided to match the user's emotional state. This emotional information is also applied to suggest actions that enhance emotional satisfaction.

[0330] Furthermore, recommendations related to environmental considerations are provided in a personalized manner that takes into account the user's emotional state. For example, a user seeking relaxation might be recommended a ski resort that offers eco-friendly services in a quiet environment, using emotional information to provide more individualized suggestions.

[0331] For example, if the emotion engine recognizes a user's voice comment saying, "I'm struggling with my snowboard turns," along with a confused expression, the server will provide detailed suggestions for improving their turning technique and even suggest simple tricks to help them maintain balance. In addition, information on eco-friendly ski resorts where the user can practice while having fun will also be provided.

[0332] Thus, the present invention aims to further enhance the winter sports experience by combining flexible feedback that responds to the user's emotions with sustainable options.

[0333] The following describes the processing flow.

[0334] Step 1:

[0335] Users can use their smartphones or cameras to film videos and videos of themselves skiing or snowboarding, and record their thoughts and challenges in audio as needed.

[0336] Step 2:

[0337] Users select video and audio data through a dedicated application and upload them to the server. The upload is performed using a secure and efficient data transfer protocol.

[0338] Step 3:

[0339] The server receives the uploaded data. Upon receipt, the server verifies the data format and performs any necessary preprocessing. This includes frame extraction and audio cleaning.

[0340] Step 4:

[0341] The server supplies video data to a multimodal AI model for technical motion analysis. Here, technical metrics such as turn precision, balance, and jump height are extracted.

[0342] Step 5:

[0343] The server analyzes the voice data to identify the user's intentions and challenges. This allows for the generation of specific technical improvement suggestions based on the user's interests.

[0344] Step 6:

[0345] The server uses an emotion engine to analyze the user's emotions from video and audio data. It uses facial expressions and tone of voice to identify the user's current emotional state (e.g., satisfaction, impatience, excitement).

[0346] Step 7:

[0347] The server customizes feedback and action suggestions based on the analysis results and the user's emotional state. The tone and level of detail of the suggestions are adjusted according to the emotional state.

[0348] Step 8:

[0349] The server generates personalized recommendations related to eco-friendly activities based on emotional information. For example, it might recommend ski resorts with a relaxing and quiet environment.

[0350] Step 9:

[0351] The device receives feedback and suggestions from the server and presents them to the user in a visual and interactive format. This allows the user to intuitively understand the feedback and plan their next steps.

[0352] (Example 2)

[0353] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0354] In winter sports, it is difficult to provide a highly satisfying experience because there is a lack of means to provide environmentally conscious recommendation information that is tailored to the emotional state of the user, in addition to improving their skills.

[0355] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0356] In this invention, the server includes means for receiving video information, analyzing technical operations using a generation AI model to identify areas for improvement, analyzing audio information, recognizing the user's emotional state to generate new action suggestions, providing environmentally conscious recommendation information that takes the emotional state into account, and recognizing the user's emotions using an emotion engine to personalize feedback. This makes it possible to provide users with personalized technical improvement feedback and recommendation information that takes sustainability into consideration.

[0357] "Video and image information" refers to data including images and actions captured by the user during winter sports activities.

[0358] A "generative AI model" refers to an artificial intelligence algorithm used to analyze received video and image information and identify areas for improvement to enhance specific technologies.

[0359] "Audio information" refers to data that includes verbal comments and impressions made by users during or after participating in winter sports.

[0360] An "emotion engine" refers to a technological element used to recognize and analyze a user's emotions from their voice and facial expressions, and to supplement the analysis results.

[0361] "Environmentally conscious recommendation information" refers to information about facilities and services that allow users to enjoy winter sports in a more sustainable way, while taking into account their emotional state.

[0362] This invention is a system for improving winter sports techniques and providing environmentally conscious recommendation information tailored to user emotions. Users use smartphones or action cameras as terminals to collect video and audio information while skiing or snowboarding. Users upload this data to a server via their terminals.

[0363] The server inputs video data into an AI model to analyze the user's technical movements. This model evaluates the angle and speed of turns, as well as balance, and identifies specific areas for improvement. The server also analyzes audio data using natural language processing technology to extract the user's self-assessment and problems. Furthermore, it uses an emotion engine to analyze the user's facial expressions and tone of voice to recognize the user's emotional state.

[0364] Based on the analysis results, the server generates feedback for technical improvements tailored to the user, and also provides environmentally conscious recommendations that match their emotional state. For example, a user who wants to relax might be recommended an eco-friendly ski resort with a quiet environment. The feedback and recommendations are sent to the user's device, where they can review them and use them to plan their next activities.

[0365] For example, if a user provides a voice comment such as, "I'm having trouble with my snowboard turns," and the emotion engine recognizes from their facial expression that they are confused, the server will provide specific advice on improving their turning technique and also suggest eco-friendly ski resorts that the user can enjoy.

[0366] Examples of prompt statements include the following:

[0367] "Please give me some specific advice on how to improve my snowboarding turning technique."

[0368] "Please recommend an eco-friendly ski resort."

[0369] This invention allows users to improve their winter sports skills while gaining an emotionally conscious and sustainable experience.

[0370] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0371] Step 1:

[0372] Users collect video and audio information

[0373] The user uses a device to record video and audio information during winter sports activities, along with comments and emotional expressions. This collects data necessary for evaluating technical actions and emotional states. The input for this step is the user's actions and audio recordings in the field, and the output is the corresponding digital data.

[0374] Step 2:

[0375] User uploads data to the server

[0376] The user sends video and audio information collected via their device to the server. The user uploads the data using a dedicated app or web interface. The input in this step is digital data stored on the user's device, and the output is the information received by the server.

[0377] Step 3:

[0378] The server analyzes the video information.

[0379] The server inputs the received video information into a generating AI model, which analyzes technical actions such as turn angle, speed, and balance. The AI ​​model uses advanced algorithms to detect detailed technical elements and evaluate the user's performance quality. The input for this step is video information, and the output is the technical evaluation result.

[0380] Step 4:

[0381] The server analyzes the audio information.

[0382] The server analyzes the audio information using natural language processing technology. It extracts self-assessments and problems from user comments, using this as foundational data to generate action suggestions. The input for this step is audio information, and the output is the results of the self-assessment and problem extraction.

[0383] Step 5:

[0384] The server recognizes emotions.

[0385] The server uses an emotion engine to recognize the user's emotional state from video and audio tones. This analysis determines the emotional situation, such as whether the user is enjoying themselves or is confused. The input for this step is video and audio tones, and the output is the result of the emotion judgment.

[0386] Step 6:

[0387] The server generates feedback and recommendation information.

[0388] The server generates user-appropriate feedback based on technical evaluations, self-assessment results, and emotional judgments. Specific advice for technical improvement and emotionally responsive, environmentally conscious recommendations are provided. The input for this step is the entire analysis result, and the output is user feedback and recommendations.

[0389] Step 7:

[0390] Users receive feedback

[0391] The user receives and confirms feedback and recommendations sent from the server through the terminal interface. Based on the received information, the user plans their next activity. The input for this step is data from the server, and the output is the information confirmed by the user.

[0392] (Application Example 2)

[0393] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0394] In winter sports, there are problems such as a lack of detailed feedback for improving technique, a lack of personalized feedback tailored to the user's emotions, and limited information on service selection that takes environmental considerations into account. This project aims to solve these technical challenges.

[0395] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0396] This invention includes a server that includes means for receiving and analyzing video data captured by the user to identify areas for technological improvement, means for analyzing audio data and generating new action suggestions based on the user's comments, an emotion engine that recognizes the user's emotions and provides personalized feedback based on the analysis results, and means for providing recommendation information related to environmental considerations based on the analysis results and recognized emotions. This enables users enjoying winter sports to receive detailed technological feedback, personalized suggestions tailored to their emotions, and information on eco-friendly options.

[0397] "Communication equipment" is a general term for information terminals used by users to transmit data to external data processing devices.

[0398] A "data processing device" refers to a computer system that analyzes video and audio data received from users and generates action suggestions and feedback.

[0399] "Video data" refers to video files that record user activity and is used for analysis to identify areas for technical improvement.

[0400] "Voice data" refers to audio files that record the user's voice and contain information useful for emotion recognition and generating action suggestions.

[0401] An "emotion engine" is a software module that recognizes a user's emotional state from voice and visual data and personalizes suggestions and feedback based on that state.

[0402] "Eco-friendly options" refer to information about sustainable services and products that take environmental considerations into account, and are included in the recommendation information provided to users.

[0403] "Feedback" is a process of providing information that indicates areas for improvement and recommended actions based on user activity, and is particularly aimed at improving users' technical skills and the quality of their experience.

[0404] This invention provides a system for offering winter sports enthusiasts an experience that balances technological advancement with environmental consideration. The system is started when a user captures video data using a smartphone and uploads it to a data processing device via a communication device.

[0405] The server analyzes the received video data using computer vision tools (e.g., OpenCV) to extract the technical characteristics of the user's movements. This information includes technical elements such as the angle of turns and balance. Furthermore, the audio data is analyzed using speech recognition tools (e.g., Google Speech-to-Text, emotion recognition API) to identify the user's comments and emotional state.

[0406] The emotion engine recognizes the user's emotions from analyzed voice and facial expression data and generates personalized feedback and action suggestions based on that information. For example, if a user comments that they are "struggling with turns," the server will provide specific action steps to improve their turning technique and also offer practice suggestions to improve their balance. Furthermore, depending on the user's emotional state, information about eco-friendly facilities will be recommended for users seeking relaxation.

[0407] An example prompt would be, "Analyze the user's skiing video and audio data, and provide personalized information based on technical evaluation and emotion recognition." This system aims to provide users with a fulfilling sports experience by combining detailed technical checks in winter sports with emotionally tailored support, offering sustainable options.

[0408] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0409] Step 1:

[0410] The system uses a device to record video and audio data captured by the user during winter sports activities. Input is the video and audio captured by the user, and output is a data file stored on the device. Specifically, the device's camera and microphone are used to record footage of activities such as riding ski lifts or skiing.

[0411] Step 2:

[0412] The terminal uploads recorded video and audio data to the server via a communication device. The input is the video and audio data files stored on the terminal, and the output is the data transferred to the server. Specifically, the terminal transmits data wirelessly via a dedicated application.

[0413] Step 3:

[0414] The server analyzes the received video data using computer vision tools (e.g., OpenCV). The input is the video data uploaded to the server, and the output is the technical features obtained through the analysis (e.g., turn angle, speed, balance). Specifically, the system processes the video data frames sequentially and calculates various metrics of the user's movement.

[0415] Step 4:

[0416] The server receives audio data, converts it to text using a speech recognition tool (e.g., Google Speech-to-Text), and then identifies the emotional state using an emotion recognition API. The input is audio data uploaded to the server, and the output is the user's comments and the recognized emotion information. Specifically, the process converts the audio data to text and extracts emotions from the content and intonation.

[0417] Step 5:

[0418] The server integrates analyzed technical features and emotional information, and uses a generative AI model to generate user-appropriate feedback and action suggestions. The input is technical features and emotional information, and the output is personalized feedback and action suggestions. Specifically, the prompt message "Analyze the user's skiing video and audio data, and provide personalized information based on technical evaluation and emotional recognition" is input to the AI ​​model, and the generated suggestions are created.

[0419] Step 6:

[0420] The server generates feedback and action suggestions, which are then sent back to the terminal and presented to the user. The input is the generated suggestions, and the output is the feedback message that the user views on the terminal. Specifically, the information is displayed to the user as text messages or video clips through a dedicated application.

[0421] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0422] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0423] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0424] [Third Embodiment]

[0425] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0426] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0427] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0428] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0429] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0430] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0431] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0432] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0433] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0434] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0435] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0436] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0437] This invention is a system designed to simultaneously improve technical skills and environmental awareness in winter sports. The system allows users to upload video and audio data captured using their individual terminals to a server, where the data is analyzed to provide personalized feedback and suggestions.

[0438] The server takes the lead in receiving video data and uses a multimodal generative AI model to analyze the user's technical movements. This analysis extracts specific performance data, such as the position of the center of gravity and the consistency of movement. Based on this, improvements and new technology proposals are made.

[0439] In addition, the server converts the audio data into text and generates specific suggestions, including new tricks and technical challenges, based on the user's comments. These suggestions are customized to the user's skill level and interests.

[0440] Furthermore, the server accesses a database related to environmental considerations and provides information recommending eco-friendly ski resorts and activities suitable for the user. This allows users to enjoy sports in a sustainable way, not just improve their skills.

[0441] The device presents these analysis results and suggestions through a user interface, displaying feedback in a visually easy-to-understand format. Based on the feedback, users can create practice plans aimed at improving their next performance or consider environmentally conscious options.

[0442] For example, if a user uploads a video of themselves snowboarding along with a voice comment saying, "I want to improve the stability of my jumps," the server will analyze data on the user's posture and speed during the jumps from the video and suggest improvements such as shifting their center of gravity forward. It will also recommend ski resorts that offer environmentally friendly facilities, providing options for future visits. The device will display this information and feedback to the user in an easy-to-understand way, allowing them to use it to improve their next activity.

[0443] The following describes the processing flow.

[0444] Step 1:

[0445] Users can use their smartphones or cameras to film videos of themselves skiing or snowboarding, and record voice comments about their performance as needed.

[0446] Step 2:

[0447] The user launches a dedicated application and selects the captured video and audio data. The user then presses the upload button to send this data to the server. The data may be compressed during transmission depending on network conditions.

[0448] Step 3:

[0449] The server receives data from the user. After receiving the data, the server performs preprocessing to convert the video and audio data into a parseable format. This includes processing each frame of the video and denoising the audio data.

[0450] Step 4:

[0451] The server analyzes the video data using a multimodal generative AI model. The AI ​​analyzes the motion frames, extracts the user's technical elements (e.g., center of gravity, turn angle, speed changes), and identifies areas for technical improvement based on this.

[0452] Step 5:

[0453] The server converts the audio data into text and analyzes the user's comments. This generates suggestions for new tricks and actions tailored to the user's interests and goals.

[0454] Step 6:

[0455] Based on the analysis results, the server recommends suitable ski resorts and activities to the user from a database of eco-friendly information. This information supports sustainable participation in sports.

[0456] Step 7:

[0457] The device receives feedback information from the server. The feedback is displayed in a visual and interactive format that is easy for the user to understand.

[0458] Step 8:

[0459] Users review feedback information to improve technology and plan new activities. They also consider the eco-friendly options offered and use that information to inform their next visit.

[0460] (Example 1)

[0461] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0462] In modern sports activities, there are limited systems that simultaneously support individual skill improvement and increased environmental awareness. In particular, in winter sports, there is a need for users to objectively evaluate their own skills, obtain concrete guidance for improvement, and acquire appropriate information to enjoy the activity in a sustainable manner. This invention aims to meet these needs by developing a system that provides feedback on users' technical actions and recommends environmentally conscious activities.

[0463] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0464] In this invention, the server includes means for receiving and analyzing video data captured by the user to identify areas for improvement in exercise technique, means for analyzing audio data to generate new action suggestions based on the user's intentions and comments, and means for providing recommendation information related to environmental protection based on the analysis results. As a result, the user can receive an objective evaluation of their technical performance and obtain specific feedback that leads to improvement, as well as receive specific information to help them choose environmentally conscious activities.

[0465] A "user" refers to a person who uses this system to improve their skills and engage in sports activities while being mindful of the environment.

[0466] "Video data" refers to video footage of sports activities filmed by users, and the analysis of this data enables technological improvements.

[0467] "Voice data" refers to audio information, including comments and instructions generated by the user, and new action suggestions are generated based on its content.

[0468] A "server" refers to a core processing unit that receives video and audio data, provides computing resources for analysis, and returns feedback to the user.

[0469] "Analysis" refers to the process of processing received data to extract useful information and generating suggestions or feedback based on that information.

[0470] A "generative AI model" refers to artificial intelligence technology that supports analytical processing, and is responsible for evaluating user behavior and generating appropriate feedback.

[0471] A "prompt sentence" refers to a sentence that is input into the AI ​​model based on the transcribed information of the audio data, and it forms the basis for feedback generation.

[0472] "Environmental protection-related recommendations" refer to information about environmentally friendly locations and activities that are useful for users to enjoy sports in a sustainable way.

[0473] "User interface" refers to the method by which users receive analysis results and suggestions from a server and view the information visualized in an easy-to-understand format.

[0474] This invention is a system that allows users to improve their winter sports activities while also making environmentally conscious choices. Specifically, users capture video and audio data using a smartphone or other camera-equipped device. This data is then uploaded to a server via an application.

[0475] The server utilizes a generative AI model to analyze the received video data. This model evaluates, for example, changes in the user's center of gravity and the accuracy of jumps frame by frame. Machine learning libraries such as PyTorch and TensorFlow are used for the analysis, enabling precise data analysis.

[0476] For audio data, speech recognition technologies such as the Google Cloud Speech-to-Text API are used to convert it to text. The converted text is then input into a generative AI model, which generates specific feedback and new suggestions based on the user's intent.

[0477] A terminal equipped with a user interface receives analysis results and suggestions from the server and displays them in a format that is easy for the user to understand. For example, it presents suggestions for improvement derived from video data and recommendations for environmentally friendly ski resorts in an organized dashboard format.

[0478] For example, if a user uploads a video of themselves snowboarding and adds a voice comment saying, "I want to improve the stability of my jumps," the system will analyze the user's posture and speed during the jump. As a result, it will generate suggestions such as, "It would be good to move your center of gravity a little further forward when you jump."

[0479] An example of a prompt message would be: "Based on the following video and audio commentary, generate technical feedback on winter sports. Please include specific improvement suggestions, taking into account the video analysis results."

[0480] By using this system, users can receive technical feedback while also making environmentally conscious choices, leading to a more fulfilling sports experience.

[0481] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0482] Step 1:

[0483] Users use smartphones or camera-equipped devices to record videos and audio commentary of winter sports. The input consists of video and audio data, which will serve as foundational data for later analysis. Specifically, users press a button to record and, if necessary, leave voice commentary.

[0484] Step 2:

[0485] The device uploads video and audio data captured by the user to a server via a dedicated app. The input consists of video and audio data from the user, and these are sent to the server as output. Specifically, the upload begins when the user taps the send button on the app screen.

[0486] Step 3:

[0487] The server receives the uploaded video data and activates the analysis module. The input is video data sent from the terminal, and the output is numerical data related to the user's technical performance. Using a generative AI model, the analysis is performed frame by frame to evaluate the user's center of gravity and consistency of movement. Specifically, the server decomposes the video into frames and evaluates each frame sequentially.

[0488] Step 4:

[0489] The server starts a speech recognition engine to convert audio data into text. The input is audio data, and the output is the user's comments in text format. Specifically, the server passes the audio data through the speech recognition engine, then converts it to text format to prepare it for analysis.

[0490] Step 5:

[0491] The server integrates the results of video analysis and audio-to-text conversion, and generates feedback using a generative AI model. The input is the analyzed video and text data, and the output is specific technical improvements and new suggestions. Based on this, suggestions such as "move the center of gravity forward when jumping" are created. Specifically, both sets of data are analyzed in an integrated manner, and prompt sentences are input into the generative AI model to obtain feedback.

[0492] Step 6:

[0493] The server provides users with environmentally friendly recommendations. Input consists of analyzed video and environmental databases, while output is information on eco-friendly activities suitable for the user. Specifically, the server retrieves information from relevant databases based on the user's location and activity data, and then makes suggestions to the user.

[0494] Step 7:

[0495] The device receives feedback and recommendation information sent from the server and displays it visually. Input consists of feedback data and recommendation information from the server, while output is displayed in a user-friendly format. Specifically, the device displays a dashboard screen, visualizing analysis results and suggestions as icons and graphs. Based on this, the user can plan their next practice session.

[0496] (Application Example 1)

[0497] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0498] Winter sports enthusiasts often have difficulty obtaining appropriate feedback for improving their technique, and they also lack sufficient information regarding post-sports nutrition. Furthermore, they tend to be indifferent to information concerning environmental considerations. Therefore, there is a need to provide a system that allows users to enjoy sports in a way that promotes both technical improvement and environmental considerations.

[0499] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0500] In this invention, the server includes means for receiving and analyzing video and image data captured by the user to identify areas for improvement in technique, means for analyzing audio data and generating new action suggestions based on the user's comments, means for providing recommendation information related to environmental considerations based on the analysis results, and means for recommending highly nutritious menus based on exercise data. This makes it possible for users to improve their sports technique while also selecting appropriate nutritional support and environmentally conscious activities after sports.

[0501] "Users" refer to individuals who participate in winter sports or those who aim to improve their skills in those sports.

[0502] "Video and image data" refers to video and image files taken by users during winter sports activities.

[0503] "Analysis" is the act of analyzing specific data and extracting useful information from it.

[0504] "Technical improvements" refer to specific changes or points of instruction needed to improve current performance.

[0505] "Audio data" refers to audio data such as comments and explanations recorded by the user.

[0506] "Action suggestions" are specific guidelines or advice that indicate what action a user should take next.

[0507] "Environmentally conscious recommendation information" refers to information that suggests eco-friendly actions and places, encouraging choices that take sustainability into consideration.

[0508] "Exercise data" refers to data that includes activity records generated when a user participates in winter sports.

[0509] "Recommending nutritious menus" means suggesting meal options that contain the necessary nutrients based on the user's activity level and physical condition.

[0510] This invention is a system that enables winter sports participants to achieve both skill improvement and environmental consideration. Users collect video and audio data during exercise using a mobile device. This data is uploaded to a server. The server analyzes this data using a generative AI model.

[0511] The server analyzes video data and extracts specific performance metrics related to the user's skills. For example, it can obtain data such as the position of the center of gravity and the consistency of movement. Audio data is converted into text based on the user's preferences and comments, and processed through a generative AI model that suggests new actions and technical challenges.

[0512] Furthermore, the server analyzes the user's activity level based on their exercise data and suggests nutritional supplements needed after exercise. This includes a process that recommends nutritious menus using local ingredients. Eco-friendly activities and locations are also suggested, enabling users to enjoy sports in a sustainable way.

[0513] As a concrete example, after a user goes skiing, the application asks, "What kind of training should I do the next day?" The system then provides new training methods and nutritional suggestions based on the user's previous activity data. An example of a prompt message would be, "Generate training suggestions and a meal plan based on the user's recent exercise data."

[0514] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0515] Step 1:

[0516] Users collect video and audio data of their winter sports activities using their mobile devices. This involves recording video using the device's camera and recording audio commentary using the microphone. This data is stored on the device and prepared for uploading to the server.

[0517] Step 2:

[0518] The terminal uploads collected video and audio data to the server. Data transfer takes place via an internet connection, and the server receives this data. The input data is originally in various formats, but it is converted to a unified format here.

[0519] Step 3:

[0520] The server analyzes the received video data. Using a generative AI model, it extracts metrics related to user performance. In this process, it analyzes data such as the center of gravity and motion patterns to identify areas for improvement. The input is video data, and the output is the analyzed performance data.

[0521] Step 4:

[0522] The server converts audio data to text and generates new action suggestions from user comments. It uses speech recognition software to convert speech to text and analyzes its content to understand the user's intent and requests. The input is audio data, and the output is action suggestions in text format.

[0523] Step 5:

[0524] The server provides environmentally conscious recommendations and nutritious menu suggestions based on analyzed performance data and action suggestions. This includes eco-friendly options and meal plans tailored to the user's activity level. Inputs are analyzed data and action suggestions, and output is provided as recommendations.

[0525] Step 6:

[0526] The terminal displays feedback received from the server to the user. Through the user interface, analysis results and suggestions are presented in an easy-to-understand format. Based on this, the user decides on their next exercise plan and meal choices. The input is feedback from the server, and the output is the notification content to the user.

[0527] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0528] This invention is a system that combines a new emotion engine to improve skills and environmental awareness in winter sports. Users upload video and audio data of themselves skiing or snowboarding to a server using their own devices. By analyzing this data, the system evaluates the user's performance in detail and provides specific feedback for skill improvement.

[0529] The server receives video data in real time and analyzes the user's movements using an AI model. This allows it to evaluate detailed technical elements such as turn angles and balance, and identify specific areas for improvement. In addition, it extracts the user's self-assessment from audio data and generates personalized action suggestions.

[0530] The newly integrated emotion engine recognizes emotions from the user's voice and facial expressions, contributing to improved analysis accuracy. Specifically, the emotion engine determines whether the user is enjoying themselves or experiencing difficulties, and adjusts the feedback and suggestions provided to match the user's emotional state. This emotional information is also applied to suggest actions that enhance emotional satisfaction.

[0531] Furthermore, recommendations related to environmental considerations are provided in a personalized manner that takes into account the user's emotional state. For example, a user seeking relaxation might be recommended a ski resort that offers eco-friendly services in a quiet environment, using emotional information to provide more individualized suggestions.

[0532] For example, if the emotion engine recognizes a user's voice comment saying, "I'm struggling with my snowboard turns," along with a confused expression, the server will provide detailed suggestions for improving their turning technique and even suggest simple tricks to help them maintain balance. In addition, information on eco-friendly ski resorts where the user can practice while having fun will also be provided.

[0533] Thus, the present invention aims to further enhance the winter sports experience by combining flexible feedback that responds to the user's emotions with sustainable options.

[0534] The following describes the processing flow.

[0535] Step 1:

[0536] Users can use their smartphones or cameras to film videos and videos of themselves skiing or snowboarding, and record their thoughts and challenges in audio as needed.

[0537] Step 2:

[0538] Users select video and audio data through a dedicated application and upload them to the server. The upload is performed using a secure and efficient data transfer protocol.

[0539] Step 3:

[0540] The server receives the uploaded data. Upon receipt, the server verifies the data format and performs any necessary preprocessing. This includes frame extraction and audio cleaning.

[0541] Step 4:

[0542] The server supplies video data to a multimodal AI model for technical motion analysis. Here, technical metrics such as turn precision, balance, and jump height are extracted.

[0543] Step 5:

[0544] The server analyzes the voice data to identify the user's intentions and challenges. This allows for the generation of specific technical improvement suggestions based on the user's interests.

[0545] Step 6:

[0546] The server uses an emotion engine to analyze the user's emotions from video and audio data. It uses facial expressions and tone of voice to identify the user's current emotional state (e.g., satisfaction, impatience, excitement).

[0547] Step 7:

[0548] The server customizes feedback and action suggestions based on the analysis results and the user's emotional state. The tone and level of detail of the suggestions are adjusted according to the emotional state.

[0549] Step 8:

[0550] The server generates personalized recommendations related to eco-friendly activities based on emotional information. For example, it might recommend ski resorts with a relaxing and quiet environment.

[0551] Step 9:

[0552] The device receives feedback and suggestions from the server and presents them to the user in a visual and interactive format. This allows the user to intuitively understand the feedback and plan their next steps.

[0553] (Example 2)

[0554] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0555] In winter sports, it is difficult to provide a highly satisfying experience because there is a lack of means to provide environmentally conscious recommendation information that is tailored to the emotional state of the user, in addition to improving their skills.

[0556] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0557] In this invention, the server includes means for receiving video information, analyzing technical operations using a generation AI model to identify areas for improvement, analyzing audio information, recognizing the user's emotional state to generate new action suggestions, providing environmentally conscious recommendation information that takes the emotional state into account, and recognizing the user's emotions using an emotion engine to personalize feedback. This makes it possible to provide users with personalized technical improvement feedback and recommendation information that takes sustainability into consideration.

[0558] "Video and image information" refers to data including images and actions captured by the user during winter sports activities.

[0559] A "generative AI model" refers to an artificial intelligence algorithm used to analyze received video and image information and identify areas for improvement to enhance specific technologies.

[0560] "Audio information" refers to data that includes verbal comments and impressions made by users during or after participating in winter sports.

[0561] An "emotion engine" refers to a technological element used to recognize and analyze a user's emotions from their voice and facial expressions, and to supplement the analysis results.

[0562] "Environmentally conscious recommendation information" refers to information about facilities and services that allow users to enjoy winter sports in a more sustainable way, while taking into account their emotional state.

[0563] This invention is a system for improving winter sports techniques and providing environmentally conscious recommendation information tailored to user emotions. Users use smartphones or action cameras as terminals to collect video and audio information while skiing or snowboarding. Users upload this data to a server via their terminals.

[0564] The server inputs video data into an AI model to analyze the user's technical movements. This model evaluates the angle and speed of turns, as well as balance, and identifies specific areas for improvement. The server also analyzes audio data using natural language processing technology to extract the user's self-assessment and problems. Furthermore, it uses an emotion engine to analyze the user's facial expressions and tone of voice to recognize the user's emotional state.

[0565] Based on the analysis results, the server generates feedback for technical improvements tailored to the user, and also provides environmentally conscious recommendations that match their emotional state. For example, a user who wants to relax might be recommended an eco-friendly ski resort with a quiet environment. The feedback and recommendations are sent to the user's device, where they can review them and use them to plan their next activities.

[0566] For example, if a user provides a voice comment such as, "I'm having trouble with my snowboard turns," and the emotion engine recognizes from their facial expression that they are confused, the server will provide specific advice on improving their turning technique and also suggest eco-friendly ski resorts that the user can enjoy.

[0567] Examples of prompt statements include the following:

[0568] "Please give me some specific advice on how to improve my snowboarding turning technique."

[0569] "Please recommend an eco-friendly ski resort."

[0570] This invention allows users to improve their winter sports skills while gaining an emotionally conscious and sustainable experience.

[0571] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0572] Step 1:

[0573] Users collect video and audio information

[0574] The user uses a device to record video and audio information during winter sports activities, along with comments and emotional expressions. This collects data necessary for evaluating technical actions and emotional states. The input for this step is the user's actions and audio recordings in the field, and the output is the corresponding digital data.

[0575] Step 2:

[0576] User uploads data to the server

[0577] The user sends video and audio information collected via their device to the server. The user uploads the data using a dedicated app or web interface. The input in this step is digital data stored on the user's device, and the output is the information received by the server.

[0578] Step 3:

[0579] The server analyzes the video information.

[0580] The server inputs the received video information into a generating AI model, which analyzes technical actions such as turn angle, speed, and balance. The AI ​​model uses advanced algorithms to detect detailed technical elements and evaluate the user's performance quality. The input for this step is video information, and the output is the technical evaluation result.

[0581] Step 4:

[0582] The server analyzes the audio information.

[0583] The server analyzes the audio information using natural language processing technology. It extracts self-assessments and problems from user comments, using this as foundational data to generate action suggestions. The input for this step is audio information, and the output is the results of the self-assessment and problem extraction.

[0584] Step 5:

[0585] The server recognizes emotions.

[0586] The server uses an emotion engine to recognize the user's emotional state from video and audio tones. This analysis determines the emotional situation, such as whether the user is enjoying themselves or is confused. The input for this step is video and audio tones, and the output is the result of the emotion judgment.

[0587] Step 6:

[0588] The server generates feedback and recommendation information.

[0589] The server generates user-appropriate feedback based on technical evaluations, self-assessment results, and emotional judgments. Specific advice for technical improvement and emotionally responsive, environmentally conscious recommendations are provided. The input for this step is the entire analysis result, and the output is user feedback and recommendations.

[0590] Step 7:

[0591] Users receive feedback

[0592] The user receives and confirms feedback and recommendations sent from the server through the terminal interface. Based on the received information, the user plans their next activity. The input for this step is data from the server, and the output is the information confirmed by the user.

[0593] (Application Example 2)

[0594] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0595] In winter sports, there are problems such as a lack of detailed feedback for improving technique, a lack of personalized feedback tailored to the user's emotions, and limited information on service selection that takes environmental considerations into account. This project aims to solve these technical challenges.

[0596] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0597] This invention includes a server that includes means for receiving and analyzing video data captured by the user to identify areas for technological improvement, means for analyzing audio data and generating new action suggestions based on the user's comments, an emotion engine that recognizes the user's emotions and provides personalized feedback based on the analysis results, and means for providing recommendation information related to environmental considerations based on the analysis results and recognized emotions. This enables users enjoying winter sports to receive detailed technological feedback, personalized suggestions tailored to their emotions, and information on eco-friendly options.

[0598] "Communication equipment" is a general term for information terminals used by users to transmit data to external data processing devices.

[0599] A "data processing device" refers to a computer system that analyzes video and audio data received from users and generates action suggestions and feedback.

[0600] "Video data" refers to video files that record user activity and is used for analysis to identify areas for technical improvement.

[0601] "Voice data" refers to audio files that record the user's voice and contain information useful for emotion recognition and generating action suggestions.

[0602] An "emotion engine" is a software module that recognizes a user's emotional state from voice and visual data and personalizes suggestions and feedback based on that state.

[0603] "Eco-friendly options" refer to information about sustainable services and products that take environmental considerations into account, and are included in the recommendation information provided to users.

[0604] "Feedback" is a process of providing information that indicates areas for improvement and recommended actions based on user activity, and is particularly aimed at improving users' technical skills and the quality of their experience.

[0605] This invention provides a system for offering winter sports enthusiasts an experience that balances technological advancement with environmental consideration. The system is started when a user captures video data using a smartphone and uploads it to a data processing device via a communication device.

[0606] The server analyzes the received video data using computer vision tools (e.g., OpenCV) to extract the technical characteristics of the user's movements. This information includes technical elements such as the angle of turns and balance. Furthermore, the audio data is analyzed using speech recognition tools (e.g., Google Speech-to-Text, emotion recognition API) to identify the user's comments and emotional state.

[0607] The emotion engine recognizes the user's emotions from analyzed voice and facial expression data and generates personalized feedback and action suggestions based on that information. For example, if a user comments that they are "struggling with turns," the server will provide specific action steps to improve their turning technique and also offer practice suggestions to improve their balance. Furthermore, depending on the user's emotional state, information about eco-friendly facilities will be recommended for users seeking relaxation.

[0608] An example prompt would be, "Analyze the user's skiing video and audio data, and provide personalized information based on technical evaluation and emotion recognition." This system aims to provide users with a fulfilling sports experience by combining detailed technical checks in winter sports with emotionally tailored support, offering sustainable options.

[0609] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0610] Step 1:

[0611] The system uses a device to record video and audio data captured by the user during winter sports activities. Input is the video and audio captured by the user, and output is a data file stored on the device. Specifically, the device's camera and microphone are used to record footage of activities such as riding ski lifts or skiing.

[0612] Step 2:

[0613] The terminal uploads recorded video and audio data to the server via a communication device. The input is the video and audio data files stored on the terminal, and the output is the data transferred to the server. Specifically, the terminal transmits data wirelessly via a dedicated application.

[0614] Step 3:

[0615] The server analyzes the received video data using computer vision tools (e.g., OpenCV). The input is the video data uploaded to the server, and the output is the technical features obtained through the analysis (e.g., turn angle, speed, balance). Specifically, the system processes the video data frames sequentially and calculates various metrics of the user's movement.

[0616] Step 4:

[0617] The server receives audio data, converts it to text using a speech recognition tool (e.g., Google Speech-to-Text), and then identifies the emotional state using an emotion recognition API. The input is audio data uploaded to the server, and the output is the user's comments and the recognized emotion information. Specifically, the process converts the audio data to text and extracts emotions from the content and intonation.

[0618] Step 5:

[0619] The server integrates analyzed technical features and emotional information, and uses a generative AI model to generate user-appropriate feedback and action suggestions. The input is technical features and emotional information, and the output is personalized feedback and action suggestions. Specifically, the prompt message "Analyze the user's skiing video and audio data, and provide personalized information based on technical evaluation and emotional recognition" is input to the AI ​​model, and the generated suggestions are created.

[0620] Step 6:

[0621] The server generates feedback and action suggestions, which are then sent back to the terminal and presented to the user. The input is the generated suggestions, and the output is the feedback message that the user views on the terminal. Specifically, the information is displayed to the user as text messages or video clips through a dedicated application.

[0622] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0623] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0624] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0625] [Fourth Embodiment]

[0626] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0627] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0628] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0629] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0630] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0631] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0632] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0633] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0634] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0635] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0636] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0637] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0638] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0639] This invention is a system designed to simultaneously improve technical skills and environmental awareness in winter sports. The system allows users to upload video and audio data captured using their individual terminals to a server, where the data is analyzed to provide personalized feedback and suggestions.

[0640] The server takes the lead in receiving video data and uses a multimodal generative AI model to analyze the user's technical movements. This analysis extracts specific performance data, such as the position of the center of gravity and the consistency of movement. Based on this, improvements and new technology proposals are made.

[0641] In addition, the server converts the audio data into text and generates specific suggestions, including new tricks and technical challenges, based on the user's comments. These suggestions are customized to the user's skill level and interests.

[0642] Furthermore, the server accesses a database related to environmental considerations and provides information recommending eco-friendly ski resorts and activities suitable for the user. This allows users to enjoy sports in a sustainable way, not just improve their skills.

[0643] The device presents these analysis results and suggestions through a user interface, displaying feedback in a visually easy-to-understand format. Based on the feedback, users can create practice plans aimed at improving their next performance or consider environmentally conscious options.

[0644] For example, if a user uploads a video of themselves snowboarding along with a voice comment saying, "I want to improve the stability of my jumps," the server will analyze data on the user's posture and speed during the jumps from the video and suggest improvements such as shifting their center of gravity forward. It will also recommend ski resorts that offer environmentally friendly facilities, providing options for future visits. The device will display this information and feedback to the user in an easy-to-understand way, allowing them to use it to improve their next activity.

[0645] The following describes the processing flow.

[0646] Step 1:

[0647] Users can use their smartphones or cameras to film videos of themselves skiing or snowboarding, and record voice comments about their performance as needed.

[0648] Step 2:

[0649] The user launches a dedicated application and selects the captured video and audio data. The user then presses the upload button to send this data to the server. The data may be compressed during transmission depending on network conditions.

[0650] Step 3:

[0651] The server receives data from the user. After receiving the data, the server performs preprocessing to convert the video and audio data into a parseable format. This includes processing each frame of the video and denoising the audio data.

[0652] Step 4:

[0653] The server analyzes the video data using a multimodal generative AI model. The AI ​​analyzes the motion frames, extracts the user's technical elements (e.g., center of gravity, turn angle, speed changes), and identifies areas for technical improvement based on this.

[0654] Step 5:

[0655] The server converts the audio data into text and analyzes the user's comments. This generates suggestions for new tricks and actions tailored to the user's interests and goals.

[0656] Step 6:

[0657] Based on the analysis results, the server recommends suitable ski resorts and activities to the user from a database of eco-friendly information. This information supports sustainable participation in sports.

[0658] Step 7:

[0659] The device receives feedback information from the server. The feedback is displayed in a visual and interactive format that is easy for the user to understand.

[0660] Step 8:

[0661] Users review feedback information to improve technology and plan new activities. They also consider the eco-friendly options offered and use that information to inform their next visit.

[0662] (Example 1)

[0663] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0664] In modern sports activities, there are limited systems that simultaneously support individual skill improvement and increased environmental awareness. In particular, in winter sports, there is a need for users to objectively evaluate their own skills, obtain concrete guidance for improvement, and acquire appropriate information to enjoy the activity in a sustainable manner. This invention aims to meet these needs by developing a system that provides feedback on users' technical actions and recommends environmentally conscious activities.

[0665] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0666] In this invention, the server includes means for receiving and analyzing video data captured by the user to identify areas for improvement in exercise technique, means for analyzing audio data to generate new action suggestions based on the user's intentions and comments, and means for providing recommendation information related to environmental protection based on the analysis results. As a result, the user can receive an objective evaluation of their technical performance and obtain specific feedback that leads to improvement, as well as receive specific information to help them choose environmentally conscious activities.

[0667] A "user" refers to a person who uses this system to improve their skills and engage in sports activities while being mindful of the environment.

[0668] "Video data" refers to video footage of sports activities filmed by users, and the analysis of this data enables technological improvements.

[0669] "Voice data" refers to audio information, including comments and instructions generated by the user, and new action suggestions are generated based on its content.

[0670] A "server" refers to a core processing unit that receives video and audio data, provides computing resources for analysis, and returns feedback to the user.

[0671] "Analysis" refers to the process of processing received data to extract useful information and generating suggestions or feedback based on that information.

[0672] A "generative AI model" refers to artificial intelligence technology that supports analytical processing, and is responsible for evaluating user behavior and generating appropriate feedback.

[0673] A "prompt sentence" refers to a sentence that is input into the AI ​​model based on the transcribed information of the audio data, and it forms the basis for feedback generation.

[0674] "Environmental protection-related recommendations" refer to information about environmentally friendly locations and activities that are useful for users to enjoy sports in a sustainable way.

[0675] "User interface" refers to the method by which users receive analysis results and suggestions from a server and view the information visualized in an easy-to-understand format.

[0676] This invention is a system that allows users to improve their winter sports activities while also making environmentally conscious choices. Specifically, users capture video and audio data using a smartphone or other camera-equipped device. This data is then uploaded to a server via an application.

[0677] The server utilizes a generative AI model to analyze the received video data. This model evaluates, for example, changes in the user's center of gravity and the accuracy of jumps frame by frame. Machine learning libraries such as PyTorch and TensorFlow are used for the analysis, enabling precise data analysis.

[0678] For audio data, speech recognition technologies such as the Google Cloud Speech-to-Text API are used to convert it to text. The converted text is then input into a generative AI model, which generates specific feedback and new suggestions based on the user's intent.

[0679] A terminal equipped with a user interface receives analysis results and suggestions from the server and displays them in a format that is easy for the user to understand. For example, it presents suggestions for improvement derived from video data and recommendations for environmentally friendly ski resorts in an organized dashboard format.

[0680] For example, if a user uploads a video of themselves snowboarding and adds a voice comment saying, "I want to improve the stability of my jumps," the system will analyze the user's posture and speed during the jump. As a result, it will generate suggestions such as, "It would be good to move your center of gravity a little further forward when you jump."

[0681] An example of a prompt message would be: "Based on the following video and audio commentary, generate technical feedback on winter sports. Please include specific improvement suggestions, taking into account the video analysis results."

[0682] By using this system, users can receive technical feedback while also making environmentally conscious choices, leading to a more fulfilling sports experience.

[0683] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0684] Step 1:

[0685] Users use smartphones or camera-equipped devices to record videos and audio commentary of winter sports. The input consists of video and audio data, which will serve as foundational data for later analysis. Specifically, users press a button to record and, if necessary, leave voice commentary.

[0686] Step 2:

[0687] The device uploads video and audio data captured by the user to a server via a dedicated app. The input consists of video and audio data from the user, and these are sent to the server as output. Specifically, the upload begins when the user taps the send button on the app screen.

[0688] Step 3:

[0689] The server receives the uploaded video data and activates the analysis module. The input is video data sent from the terminal, and the output is numerical data related to the user's technical performance. Using a generative AI model, the analysis is performed frame by frame to evaluate the user's center of gravity and consistency of movement. Specifically, the server decomposes the video into frames and evaluates each frame sequentially.

[0690] Step 4:

[0691] The server starts a speech recognition engine to convert audio data into text. The input is audio data, and the output is the user's comments in text format. Specifically, the server passes the audio data through the speech recognition engine, then converts it to text format to prepare it for analysis.

[0692] Step 5:

[0693] The server integrates the results of video analysis and audio-to-text conversion, and generates feedback using a generative AI model. The input is the analyzed video and text data, and the output is specific technical improvements and new suggestions. Based on this, suggestions such as "move the center of gravity forward when jumping" are created. Specifically, both sets of data are analyzed in an integrated manner, and prompt sentences are input into the generative AI model to obtain feedback.

[0694] Step 6:

[0695] The server provides users with environmentally friendly recommendations. Input consists of analyzed video and environmental databases, while output is information on eco-friendly activities suitable for the user. Specifically, the server retrieves information from relevant databases based on the user's location and activity data, and then makes suggestions to the user.

[0696] Step 7:

[0697] The device receives feedback and recommendation information sent from the server and displays it visually. Input consists of feedback data and recommendation information from the server, while output is displayed in a user-friendly format. Specifically, the device displays a dashboard screen, visualizing analysis results and suggestions as icons and graphs. Based on this, the user can plan their next practice session.

[0698] (Application Example 1)

[0699] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0700] Winter sports enthusiasts often have difficulty obtaining appropriate feedback for improving their technique, and they also lack sufficient information regarding post-sports nutrition. Furthermore, they tend to be indifferent to information concerning environmental considerations. Therefore, there is a need to provide a system that allows users to enjoy sports in a way that promotes both technical improvement and environmental considerations.

[0701] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0702] In this invention, the server includes means for receiving and analyzing video and image data captured by the user to identify areas for improvement in technique, means for analyzing audio data and generating new action suggestions based on the user's comments, means for providing recommendation information related to environmental considerations based on the analysis results, and means for recommending highly nutritious menus based on exercise data. This makes it possible for users to improve their sports technique while also selecting appropriate nutritional support and environmentally conscious activities after sports.

[0703] "Users" refer to individuals who participate in winter sports or those who aim to improve their skills in those sports.

[0704] "Video and image data" refers to video and image files taken by users during winter sports activities.

[0705] "Analysis" is the act of analyzing specific data and extracting useful information from it.

[0706] "Technical improvements" refer to specific changes or points of instruction needed to improve current performance.

[0707] "Audio data" refers to audio data such as comments and explanations recorded by the user.

[0708] "Action suggestions" are specific guidelines or advice that indicate what action a user should take next.

[0709] "Environmentally conscious recommendation information" refers to information that suggests eco-friendly actions and places, encouraging choices that take sustainability into consideration.

[0710] "Exercise data" refers to data that includes activity records generated when a user participates in winter sports.

[0711] "Recommending nutritious menus" means suggesting meal options that contain the necessary nutrients based on the user's activity level and physical condition.

[0712] This invention is a system that enables winter sports participants to achieve both skill improvement and environmental consideration. Users collect video and audio data during exercise using a mobile device. This data is uploaded to a server. The server analyzes this data using a generative AI model.

[0713] The server analyzes video data and extracts specific performance metrics related to the user's skills. For example, it can obtain data such as the position of the center of gravity and the consistency of movement. Audio data is converted into text based on the user's preferences and comments, and processed through a generative AI model that suggests new actions and technical challenges.

[0714] Furthermore, the server analyzes the user's activity level based on their exercise data and suggests nutritional supplements needed after exercise. This includes a process that recommends nutritious menus using local ingredients. Eco-friendly activities and locations are also suggested, enabling users to enjoy sports in a sustainable way.

[0715] As a concrete example, after a user goes skiing, the application asks, "What kind of training should I do the next day?" The system then provides new training methods and nutritional suggestions based on the user's previous activity data. An example of a prompt message would be, "Generate training suggestions and a meal plan based on the user's recent exercise data."

[0716] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0717] Step 1:

[0718] Users collect video and audio data of their winter sports activities using their mobile devices. This involves recording video using the device's camera and recording audio commentary using the microphone. This data is stored on the device and prepared for uploading to the server.

[0719] Step 2:

[0720] The terminal uploads collected video and audio data to the server. Data transfer takes place via an internet connection, and the server receives this data. The input data is originally in various formats, but it is converted to a unified format here.

[0721] Step 3:

[0722] The server analyzes the received video data. Using a generative AI model, it extracts metrics related to user performance. In this process, it analyzes data such as the center of gravity and motion patterns to identify areas for improvement. The input is video data, and the output is the analyzed performance data.

[0723] Step 4:

[0724] The server converts audio data to text and generates new action suggestions from user comments. It uses speech recognition software to convert speech to text and analyzes its content to understand the user's intent and requests. The input is audio data, and the output is action suggestions in text format.

[0725] Step 5:

[0726] The server provides environmentally conscious recommendations and nutritious menu suggestions based on analyzed performance data and action suggestions. This includes eco-friendly options and meal plans tailored to the user's activity level. Inputs are analyzed data and action suggestions, and output is provided as recommendations.

[0727] Step 6:

[0728] The terminal displays feedback received from the server to the user. Through the user interface, analysis results and suggestions are presented in an easy-to-understand format. Based on this, the user decides on their next exercise plan and meal choices. The input is feedback from the server, and the output is the notification content to the user.

[0729] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0730] This invention is a system that combines a new emotion engine to improve skills and environmental awareness in winter sports. Users upload video and audio data of themselves skiing or snowboarding to a server using their own devices. By analyzing this data, the system evaluates the user's performance in detail and provides specific feedback for skill improvement.

[0731] The server receives video data in real time and analyzes the user's movements using an AI model. This allows it to evaluate detailed technical elements such as turn angles and balance, and identify specific areas for improvement. In addition, it extracts the user's self-assessment from audio data and generates personalized action suggestions.

[0732] The newly integrated emotion engine recognizes emotions from the user's voice and facial expressions, contributing to improved analysis accuracy. Specifically, the emotion engine determines whether the user is enjoying themselves or experiencing difficulties, and adjusts the feedback and suggestions provided to match the user's emotional state. This emotional information is also applied to suggest actions that enhance emotional satisfaction.

[0733] Furthermore, recommendations related to environmental considerations are provided in a personalized manner that takes into account the user's emotional state. For example, a user seeking relaxation might be recommended a ski resort that offers eco-friendly services in a quiet environment, using emotional information to provide more individualized suggestions.

[0734] For example, if the emotion engine recognizes a user's voice comment saying, "I'm struggling with my snowboard turns," along with a confused expression, the server will provide detailed suggestions for improving their turning technique and even suggest simple tricks to help them maintain balance. In addition, information on eco-friendly ski resorts where the user can practice while having fun will also be provided.

[0735] Thus, the present invention aims to further enhance the winter sports experience by combining flexible feedback that responds to the user's emotions with sustainable options.

[0736] The following describes the processing flow.

[0737] Step 1:

[0738] Users can use their smartphones or cameras to film videos and videos of themselves skiing or snowboarding, and record their thoughts and challenges in audio as needed.

[0739] Step 2:

[0740] Users select video and audio data through a dedicated application and upload them to the server. The upload is performed using a secure and efficient data transfer protocol.

[0741] Step 3:

[0742] The server receives the uploaded data. Upon receipt, the server verifies the data format and performs any necessary preprocessing. This includes frame extraction and audio cleaning.

[0743] Step 4:

[0744] The server supplies video data to a multimodal AI model for technical motion analysis. Here, technical metrics such as turn precision, balance, and jump height are extracted.

[0745] Step 5:

[0746] The server analyzes the voice data to identify the user's intentions and challenges. This allows for the generation of specific technical improvement suggestions based on the user's interests.

[0747] Step 6:

[0748] The server uses an emotion engine to analyze the user's emotions from video and audio data. It uses facial expressions and tone of voice to identify the user's current emotional state (e.g., satisfaction, impatience, excitement).

[0749] Step 7:

[0750] The server customizes feedback and action suggestions based on the analysis results and the user's emotional state. The tone and level of detail of the suggestions are adjusted according to the emotional state.

[0751] Step 8:

[0752] The server generates personalized recommendations related to eco-friendly activities based on emotional information. For example, it might recommend ski resorts with a relaxing and quiet environment.

[0753] Step 9:

[0754] The device receives feedback and suggestions from the server and presents them to the user in a visual and interactive format. This allows the user to intuitively understand the feedback and plan their next steps.

[0755] (Example 2)

[0756] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0757] In winter sports, it is difficult to provide a highly satisfying experience because there is a lack of means to provide environmentally conscious recommendation information that is tailored to the emotional state of the user, in addition to improving their skills.

[0758] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0759] In this invention, the server includes means for receiving video information, analyzing technical operations using a generation AI model to identify areas for improvement, analyzing audio information, recognizing the user's emotional state to generate new action suggestions, providing environmentally conscious recommendation information that takes the emotional state into account, and recognizing the user's emotions using an emotion engine to personalize feedback. This makes it possible to provide users with personalized technical improvement feedback and recommendation information that takes sustainability into consideration.

[0760] "Video and image information" refers to data including images and actions captured by the user during winter sports activities.

[0761] A "generative AI model" refers to an artificial intelligence algorithm used to analyze received video and image information and identify areas for improvement to enhance specific technologies.

[0762] "Audio information" refers to data that includes verbal comments and impressions made by users during or after participating in winter sports.

[0763] An "emotion engine" refers to a technological element used to recognize and analyze a user's emotions from their voice and facial expressions, and to supplement the analysis results.

[0764] "Environmentally conscious recommendation information" refers to information about facilities and services that allow users to enjoy winter sports in a more sustainable way, while taking into account their emotional state.

[0765] This invention is a system for improving winter sports techniques and providing environmentally conscious recommendation information tailored to user emotions. Users use smartphones or action cameras as terminals to collect video and audio information while skiing or snowboarding. Users upload this data to a server via their terminals.

[0766] The server inputs video data into an AI model to analyze the user's technical movements. This model evaluates the angle and speed of turns, as well as balance, and identifies specific areas for improvement. The server also analyzes audio data using natural language processing technology to extract the user's self-assessment and problems. Furthermore, it uses an emotion engine to analyze the user's facial expressions and tone of voice to recognize the user's emotional state.

[0767] Based on the analysis results, the server generates feedback for technical improvements tailored to the user, and also provides environmentally conscious recommendations that match their emotional state. For example, a user who wants to relax might be recommended an eco-friendly ski resort with a quiet environment. The feedback and recommendations are sent to the user's device, where they can review them and use them to plan their next activities.

[0768] For example, if a user provides a voice comment such as, "I'm having trouble with my snowboard turns," and the emotion engine recognizes from their facial expression that they are confused, the server will provide specific advice on improving their turning technique and also suggest eco-friendly ski resorts that the user can enjoy.

[0769] Examples of prompt statements include the following:

[0770] "Please give me some specific advice on how to improve my snowboarding turning technique."

[0771] "Please recommend an eco-friendly ski resort."

[0772] This invention allows users to improve their winter sports skills while gaining an emotionally conscious and sustainable experience.

[0773] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0774] Step 1:

[0775] Users collect video and audio information

[0776] The user uses a device to record video and audio information during winter sports activities, along with comments and emotional expressions. This collects data necessary for evaluating technical actions and emotional states. The input for this step is the user's actions and audio recordings in the field, and the output is the corresponding digital data.

[0777] Step 2:

[0778] User uploads data to the server

[0779] The user sends video and audio information collected via their device to the server. The user uploads the data using a dedicated app or web interface. The input in this step is digital data stored on the user's device, and the output is the information received by the server.

[0780] Step 3:

[0781] The server analyzes the video information.

[0782] The server inputs the received video information into a generating AI model, which analyzes technical actions such as turn angle, speed, and balance. The AI ​​model uses advanced algorithms to detect detailed technical elements and evaluate the user's performance quality. The input for this step is video information, and the output is the technical evaluation result.

[0783] Step 4:

[0784] The server analyzes the audio information.

[0785] The server analyzes the audio information using natural language processing technology. It extracts self-assessments and problems from user comments, using this as foundational data to generate action suggestions. The input for this step is audio information, and the output is the results of the self-assessment and problem extraction.

[0786] Step 5:

[0787] The server recognizes emotions.

[0788] The server uses an emotion engine to recognize the user's emotional state from video and audio tones. This analysis determines the emotional situation, such as whether the user is enjoying themselves or is confused. The input for this step is video and audio tones, and the output is the result of the emotion judgment.

[0789] Step 6:

[0790] The server generates feedback and recommendation information.

[0791] The server generates user-appropriate feedback based on technical evaluations, self-assessment results, and emotional judgments. Specific advice for technical improvement and emotionally responsive, environmentally conscious recommendations are provided. The input for this step is the entire analysis result, and the output is user feedback and recommendations.

[0792] Step 7:

[0793] Users receive feedback

[0794] The user receives and confirms feedback and recommendations sent from the server through the terminal interface. Based on the received information, the user plans their next activity. The input for this step is data from the server, and the output is the information confirmed by the user.

[0795] (Application Example 2)

[0796] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0797] In winter sports, there are problems such as a lack of detailed feedback for improving technique, a lack of personalized feedback tailored to the user's emotions, and limited information on service selection that takes environmental considerations into account. This project aims to solve these technical challenges.

[0798] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0799] This invention includes a server that includes means for receiving and analyzing video data captured by the user to identify areas for technological improvement, means for analyzing audio data and generating new action suggestions based on the user's comments, an emotion engine that recognizes the user's emotions and provides personalized feedback based on the analysis results, and means for providing recommendation information related to environmental considerations based on the analysis results and recognized emotions. This enables users enjoying winter sports to receive detailed technological feedback, personalized suggestions tailored to their emotions, and information on eco-friendly options.

[0800] "Communication equipment" is a general term for information terminals used by users to transmit data to external data processing devices.

[0801] A "data processing device" refers to a computer system that analyzes video and audio data received from users and generates action suggestions and feedback.

[0802] "Video data" refers to video files that record user activity and is used for analysis to identify areas for technical improvement.

[0803] "Voice data" refers to audio files that record the user's voice and contain information useful for emotion recognition and generating action suggestions.

[0804] An "emotion engine" is a software module that recognizes a user's emotional state from voice and visual data and personalizes suggestions and feedback based on that state.

[0805] "Eco-friendly options" refer to information about sustainable services and products that take environmental considerations into account, and are included in the recommendation information provided to users.

[0806] "Feedback" is a process of providing information that indicates areas for improvement and recommended actions based on user activity, and is particularly aimed at improving users' technical skills and the quality of their experience.

[0807] This invention provides a system for offering winter sports enthusiasts an experience that balances technological advancement with environmental consideration. The system is started when a user captures video data using a smartphone and uploads it to a data processing device via a communication device.

[0808] The server analyzes the received video data using computer vision tools (e.g., OpenCV) to extract the technical characteristics of the user's movements. This information includes technical elements such as the angle of turns and balance. Furthermore, the audio data is analyzed using speech recognition tools (e.g., Google Speech-to-Text, emotion recognition API) to identify the user's comments and emotional state.

[0809] The emotion engine recognizes the user's emotions from analyzed voice and facial expression data and generates personalized feedback and action suggestions based on that information. For example, if a user comments that they are "struggling with turns," the server will provide specific action steps to improve their turning technique and also offer practice suggestions to improve their balance. Furthermore, depending on the user's emotional state, information about eco-friendly facilities will be recommended for users seeking relaxation.

[0810] An example prompt would be, "Analyze the user's skiing video and audio data, and provide personalized information based on technical evaluation and emotion recognition." This system aims to provide users with a fulfilling sports experience by combining detailed technical checks in winter sports with emotionally tailored support, offering sustainable options.

[0811] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0812] Step 1:

[0813] The system uses a device to record video and audio data captured by the user during winter sports activities. Input is the video and audio captured by the user, and output is a data file stored on the device. Specifically, the device's camera and microphone are used to record footage of activities such as riding ski lifts or skiing.

[0814] Step 2:

[0815] The terminal uploads recorded video and audio data to the server via a communication device. The input is the video and audio data files stored on the terminal, and the output is the data transferred to the server. Specifically, the terminal transmits data wirelessly via a dedicated application.

[0816] Step 3:

[0817] The server analyzes the received video data using computer vision tools (e.g., OpenCV). The input is the video data uploaded to the server, and the output is the technical features obtained through the analysis (e.g., turn angle, speed, balance). Specifically, the system processes the video data frames sequentially and calculates various metrics of the user's movement.

[0818] Step 4:

[0819] The server receives audio data, converts it to text using a speech recognition tool (e.g., Google Speech-to-Text), and then identifies the emotional state using an emotion recognition API. The input is audio data uploaded to the server, and the output is the user's comments and the recognized emotion information. Specifically, the process converts the audio data to text and extracts emotions from the content and intonation.

[0820] Step 5:

[0821] The server integrates analyzed technical features and emotional information, and uses a generative AI model to generate user-appropriate feedback and action suggestions. The input is technical features and emotional information, and the output is personalized feedback and action suggestions. Specifically, the prompt message "Analyze the user's skiing video and audio data, and provide personalized information based on technical evaluation and emotional recognition" is input to the AI ​​model, and the generated suggestions are created.

[0822] Step 6:

[0823] The server generates feedback and action suggestions, which are then sent back to the terminal and presented to the user. The input is the generated suggestions, and the output is the feedback message that the user views on the terminal. Specifically, the information is displayed to the user as text messages or video clips through a dedicated application.

[0824] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0825] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0826] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0827] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0828] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0829] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0830] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0831] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0832] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0833] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0834] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0835] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0836] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0837] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0838] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0839] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0840] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0841] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0842] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0843] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0844] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0845] The following is further disclosed regarding the embodiments described above.

[0846] (Claim 1)

[0847] A means of identifying areas for technological improvement by receiving and analyzing video data captured by users,

[0848] A means of analyzing voice data and generating new action suggestions based on user comments,

[0849] A means of providing recommendation information related to environmental considerations based on the analysis results,

[0850] A system that includes this.

[0851] (Claim 2)

[0852] The system according to claim 1, further comprising means for a user to upload video data and audio data to a server.

[0853] (Claim 3)

[0854] The system according to claim 1, comprising means for preprocessing data received by the server and converting it into an analyzable format.

[0855] "Example 1"

[0856] (Claim 1)

[0857] A means of identifying areas for improvement in movement techniques by receiving and analyzing video data captured by users,

[0858] A means of analyzing voice data and generating new action suggestions based on user intent and comments,

[0859] A means of providing recommendation information related to environmental protection based on the analysis results,

[0860] A means for analyzing motion data frame by frame and evaluating the consistency of motion,

[0861] A means of converting audio data into text,

[0862] A means of visually presenting analysis results through a user interface,

[0863] A system that includes this.

[0864] (Claim 2)

[0865] The system according to claim 1, further comprising means for a user to upload video image data and audio data to a data processing device.

[0866] (Claim 3)

[0867] The system according to claim 1, comprising means for a data processing device to preprocess received data and convert it into an analyzable format.

[0868] "Application Example 1"

[0869] (Claim 1)

[0870] A means of identifying areas for technological improvement by receiving and analyzing video data captured by users,

[0871] A means of analyzing voice data and generating new action suggestions based on user comments,

[0872] A means of providing recommendation information related to environmental considerations based on the analysis results,

[0873] A method for recommending highly nutritious meals based on exercise data,

[0874] A system that includes this.

[0875] (Claim 2)

[0876] The system according to claim 1, further comprising means for a user to upload video data and audio data to a server.

[0877] (Claim 3)

[0878] The system according to claim 1, comprising means for preprocessing data received by the server and converting it into an analyzable format.

[0879] "Example 2 of combining an emotion engine"

[0880] (Claim 1)

[0881] A means of receiving video information, analyzing technical operations using a generated AI model, and identifying areas for improvement,

[0882] A means of analyzing voice information, recognizing the user's emotional state, and generating new action suggestions,

[0883] A means of providing environmentally conscious recommendation information that takes emotional states into account,

[0884] A means of recognizing user emotions using an emotion engine and personalizing feedback,

[0885] A system that includes this.

[0886] (Claim 2)

[0887] The system according to claim 1, further comprising means for a user to transmit video information and audio information to an information processing device.

[0888] (Claim 3)

[0889] The system according to claim 1, comprising means for preprocessing information received by an information processing device and converting it into an analyzable format.

[0890] "Application example 2 when combining with an emotional engine"

[0891] (Claim 1)

[0892] A means of identifying areas for technological improvement by receiving and analyzing video data captured by users,

[0893] A means of analyzing voice data and generating new action suggestions based on user comments,

[0894] An emotion engine that recognizes user emotions and provides personalized feedback based on the analysis results,

[0895] A means of providing recommendation information related to environmental considerations based on analysis results and perceived emotions,

[0896] A system that includes this.

[0897] (Claim 2)

[0898] The system according to claim 1, further comprising means for a user to upload video image data and audio data to a data processing device using a communication device.

[0899] (Claim 3)

[0900] The system according to claim 1, comprising means for a data processing device to preprocess received data and convert it into an analyzable format. [Explanation of Symbols]

[0901] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of identifying areas for technological improvement by receiving and analyzing video data captured by users, A means of analyzing voice data and generating new action suggestions based on user comments, A means of providing recommendation information related to environmental considerations based on the analysis results, A system that includes this.

2. The system according to claim 1, further comprising means for a user to upload video data and audio data to a server.

3. The system according to claim 1, comprising means for preprocessing data received by the server and converting it into an analyzable format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A