system
The system objectively evaluates sports players' aptitudes by analyzing movement characteristics and responses, suggesting optimal positions using machine learning, thereby enhancing performance and potential.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-25
AI Technical Summary
Conventional methods for determining a sports player's optimal position are subjective and lack a general evaluation system, leading to biased assessments and underutilization of a player's potential.
A system that extracts movement characteristics from video footage, integrates user responses to sports-related questions, and uses a machine learning model to objectively evaluate and suggest the optimal sports position based on past athlete data.
Provides an objective and efficient evaluation of a player's aptitudes, enabling reliable identification of appropriate positions and maximizing their potential.
Smart Images

Figure 2026085694000001_ABST
Abstract
Description
Technical Field
[0004] , , , ,
[0005] , , , , , ,
[0001] The technology disclosed herein relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional means for a sports player to discover their suitable position are very limited and often rely on subjective judgment. For this reason, it has been difficult for players to find the optimal position to maximize their original potential, and the problem has been that appropriate evaluation cannot be obtained. Furthermore, since there is no general evaluation system applicable to other sports, there is a problem that the verification of a player's suitability is biased depending on the sport.
Means for Solving the Problems
[0005] This invention includes a feature extraction means for extracting movement characteristics from video footage of a user's movements, an answer receiving means for presenting sports-related questions to the user and receiving their answers, and an evaluation means for integrating this information to scientifically assess the user's aptitude. By adding a suggestion means that proposes the optimal sports position using a machine learning model while comparing it with a past database, this invention provides a system that enables objective and efficient evaluation of a player's new talents and aptitudes, and reliably identifies appropriate positions in various sports.
[0006] A "recording device" is a device used to record the user's actions and is a device that generates video signals.
[0007] A "video signal" is digital or analog video information acquired from a camera or recording device.
[0008] A "signal acquisition means" is a component that has the function of receiving and processing video signals from a camera.
[0009] A "feature extraction means" is a device or system that processes and extracts the characteristics of a user's actions from an input video signal.
[0010] A "question" is a question about sports presented to the user, and its purpose is to understand the player's characteristics and skills through their answers.
[0011] A "response receiving system" is a system equipped with the function of receiving and recording response data to questions from users.
[0012] "Evaluation means" refers to a device or program for evaluating a user's athletic ability based on feature data obtained by a feature extraction means and response data collected by a response reception means.
[0013] A "suggestion method" is a system that suggests the optimal sports position to the user based on evaluation results obtained using an evaluation method.
[0014] A "machine learning model" is a computer program or algorithm that learns patterns from input data and uses them to make predictions or classifications.
[0015] A "database" is a collection of information, including data on past athletes, that has been accumulated for evaluation and comparison purposes. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a processor with a reference number (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units.An example of an arithmetic unit includes a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a RAM (Random Access Memory) with a reference number is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This system aims to support users in performing sports from the optimal position, and is a sports aptitude evaluation system that utilizes video recognition technology and machine learning. The embodiments for carrying out the present invention will be described in detail below.
[0038] The user first records their sports play in video format using a recording device. This recorded video is then uploaded from the terminal to a server via the network. The server processes the received video and performs motion analysis. Specifically, the server divides the video into frames and uses image recognition technology to extract the motion of each frame as numerical data. This feature data is useful for detailed analysis of the user's movements, physique, form, etc.
[0039] Simultaneously, users are required to answer questions via their device regarding their sports experience, playing style, and preferred movements. The device then sends these answers to a server. This question-and-answer data integrates the user's subjective information with objective data.
[0040] The server combines feature data and question-answer data, and uses a machine learning model to evaluate the most suitable position for the user based on this information. It compares this information with data on athletes obtained from past databases to assess suitability. This makes it possible to recommend the position best suited to the user's movement style and characteristics.
[0041] The recommended results are sent from the server to the user's device, which then displays them to the user. The user can review the suggested results and use them to practice and train for the newly evaluated position. For example, if the server recommends outfielder as the optimal position based on the user's pitching motion and running ability, the user can use that feedback to refine their suitability as an outfielder.
[0042] In this way, this system maximizes the user's potential and supports the improvement and optimization of their playing style in sports.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The user records their sports play with a recording device and saves the video file to their device. Using the device's interface, they upload the recorded video to the server. The server receives the uploaded video file and saves it to its storage.
[0046] Step 2:
[0047] The server divides the saved video into frames and applies an image recognition algorithm to each frame to extract motion characteristics. During this process, the server collects information such as the user's posture, movement speed, and angle as numerical data. The extracted data is temporarily stored.
[0048] Step 3:
[0049] The user is asked to answer questions about their sports experience. The device displays a question form, and the user enters information about their experience, playing style, and preferred movements. The device then sends these answers to the server.
[0050] Step 4:
[0051] The server receives image recognition feature data extracted from the frames and question data answered by the user, and integrates this data. The server then performs preprocessing to prepare the data for use in machine learning models.
[0052] Step 5:
[0053] The server inputs the prepared data into a machine learning model and begins the process of evaluating the user's suitability. The model analyzes influential features and performs calculations to recommend the optimal position by comparing them with past data.
[0054] Step 6:
[0055] The server sends the user the optimal sports position derived from machine learning. The user's device receives this information and displays the results in a visually appealing format. Based on this information, the user adjusts their training and plans how to adapt to the suggested position.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] In today's sports environment, quickly and accurately determining the optimal playing position based on a player's characteristics is a challenging task. Traditional methods often rely on subjective evaluations and lack objective, data-driven analysis, which can prevent players from fully realizing their potential.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes data acquisition means, feature analysis means, and information gathering means. This makes it possible to integrate and evaluate the user's movement characteristics and subjective information, and objectively recommend the optimal movement position.
[0061] "Data acquisition means" refers to a function that inputs video data acquired from a camera that records the user's actions.
[0062] A "feature analysis method" is a function that divides video data into frames and extracts the actions of each frame as numerical data.
[0063] An "information gathering tool" is a function that presents questions to the user and receives the answers in digital format.
[0064] The "analysis means" is a function that integrates numerical data extracted by the feature analysis means with response data received by the information collection means to evaluate the user's characteristics.
[0065] "Recommendation method" refers to a function that recommends the optimal exercise position to the user based on the evaluation results from the analysis method.
[0066] This invention is a system that analyzes the user's movement characteristics based on data and proposes the optimal sports position. Specific embodiments of this system are shown below.
[0067] System configuration:
[0068] The system primarily consists of a server, terminals, and user operations. Users record their sports play in video format using recording devices (e.g., smartphones or dedicated cameras). Users upload these recorded videos to the server via their terminals. The server divides the received video into frames and uses image recognition technology as a feature analysis tool to extract and analyze the actions in each frame as numerical data. In particular, the use of deep learning libraries and computer vision models is conceivable.
[0069] Data collection:
[0070] Users answer questions about their sports experience and preferred playing style through a terminal. This information is entered into the terminal in digital format and then transmitted to a server. The server aggregates a large amount of question-and-answer data through this information collection method.
[0071] Data integration and analysis:
[0072] The server integrates numerical data extracted by feature analysis tools with question response data collected by information gathering tools. Based on the resulting dataset, machine learning techniques are used as an analytical tool to objectively evaluate the user's characteristics. The generative AI model compares the user with a database of past athletes to recommend the most suitable sports position for the user.
[0073] Presentation of results:
[0074] Based on the evaluation results, the server utilizes recommended methods to present the optimal sports position to the user on the device. The device can then display feedback to the user and provide specific training guidelines.
[0075] Specific example:
[0076] For example, if the server analyzes a user's video data and recommends that "outfielder" is the best position based on their pitching motion and excellent running ability, the user can then train based on this evaluation.
[0077] Example prompts for generative AI models:
[0078] "Evaluate the user's pitching and running abilities based on their video data, and recommend a suitable sports position."
[0079] "Consider your experience and playing style, and select the position that best suits you."
[0080] This invention provides a precise approach to maximizing sports performance.
[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0082] Step 1:
[0083] The user records their sports play in video format using a recording device. The input is the user's actions, and video data is generated as output. Specifically, the user uses a smartphone to film their soccer dribbling movements.
[0084] Step 2:
[0085] The device uploads video data to the server. The input is video data, and the output is data transfer to the server via the network. For example, pressing the "upload" button in the dedicated application sends the video to the server.
[0086] Step 3:
[0087] The server divides the received video into frames. The input is the video data transferred to the server, and the output is a dataset of image frames. Specifically, the server divides the video into 30 frames per second and generates a still image for each frame.
[0088] Step 4:
[0089] The server analyzes each frame of the video and uses image recognition technology as a feature analysis tool to extract motion characteristics. The input is image frame data, and the output is numerical data representing the motion of each frame. Specifically, the server detects things like the position of the feet and the angle of the body, and quantifies these.
[0090] Step 5:
[0091] Users answer questions about their sports experience and preferred playing style through their device. The input is user-provided data, and structured response information is generated as output. Specifically, users answer questions in a survey format such as, "What are your preferred movements?"
[0092] Step 6:
[0093] The terminal sends the user's response data to the server. The input is the user's response information, and the output is the transfer to the server. Specifically, the user presses the "Send" button on the terminal, and the data is sent to the server.
[0094] Step 7:
[0095] The server integrates numerical data extracted by the feature analysis method with response data. The input consists of quantified behavioral data and response data, and the output is an integrated dataset. This data integration enables comprehensive analysis.
[0096] Step 8:
[0097] The server uses integrated data to perform analysis with a generative AI model and evaluate the user's characteristics. The input is the integrated dataset, and the output is a proposal as an evaluation result. Specifically, the server analyzes behavioral patterns using a machine learning algorithm and recommends the optimal position.
[0098] Step 9:
[0099] The server sends the results to the terminal, which then displays the optimal position to the user. The input is the evaluation result sent from the server, and the output is the feedback displayed to the user. The user can then use this feedback to train.
[0100] (Application Example 1)
[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0102] There is a need to improve the efficiency of robots that perform multiple tasks in factories. However, currently, it is difficult to understand the optimal work position and role for each robot, making it challenging to place them in the right place. This leads to decreased productivity and wasted resources.
[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0104] In this invention, the server includes data acquisition means for inputting video signals obtained from a camera for recording user movements, data analysis means for extracting characteristics of the user's movements based on the video signals, and information receiving means for presenting questions to the user and receiving their answers. This makes it possible to recommend the optimal role and position for each robot to perform its work most efficiently.
[0105] The "data acquisition means" is responsible for acquiring video signals from devices that record the movements of users or robots, and transmitting them to a server for analysis.
[0106] A "data analysis means" is a device that uses a visual data recognition algorithm to extract motion characteristics from acquired video signals and analyzes them as numerical data.
[0107] An "information reception method" is a system that presents questions to users or administrators, receives their answers, and transmits them to a server as data for use in suitability assessments.
[0108] "Suitability evaluation means" refers to a system that integrates feature data obtained from data analysis means and information data acquired from information reception means to evaluate the most suitable work position and role for each robot.
[0109] The "role proposal method" is a process that proposes the most suitable work role for a robot based on the evaluation results derived from the aptitude evaluation method, and supports improvements in placement and work content.
[0110] This system is designed to optimize the work efficiency of robots in factories. First, cameras are attached to the robots in the factory, and these cameras continuously record the robots' movements. This serves as a means of data acquisition.
[0111] The server uses video recognition technologies such as OpenCV to divide the acquired video signal into frames and processes the motion of each frame using data analysis tools. This process extracts the robot's motion characteristics and external features as numerical data. Furthermore, the server uses machine learning frameworks such as TENSORFLOW® and PyTorch to analyze the data.
[0112] As a means of receiving information, administrators are provided with tablets or PCs, on which questions from the server are presented. Administrators answer these questions, and the content of their answers is sent to the server. The server uses these answers to evaluate the robot's suitability using an aptitude evaluation system.
[0113] As a result, the optimal robot work position and role are suggested by the role suggestion system. This information can be viewed on the factory's management system.
[0114] For example, if a suitability assessment recommends that a particular robot be specialized for transport tasks in order to improve the efficiency of assembly and transport operations, the factory manager can change the robot's placement based on that suggestion.
[0115] An example of a prompt in a generative AI model is, "Based on the operation data of the factory robot, please suggest the most efficient work role."
[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0117] Step 1:
[0118] A factory robot equipped with a camera records its own movements as video during work. The input is real-time video data. This video data is later uploaded to a server as a data acquisition method.
[0119] Step 2:
[0120] The server divides the uploaded video data into frames. The input is the video data, and the output is a collection of still images divided into each frame. The server uses these divided frames to prepare for extracting the operating characteristics using OpenCV.
[0121] Step 3:
[0122] The server uses OpenCV as a data analysis tool to extract numerical data representing the robot's motion characteristics from each frame. The input is the divided frames, and the output is numerical data showing the motion characteristics. The server then uses this numerical data for subsequent analysis.
[0123] Step 4:
[0124] The server analyzes numerical data using machine learning frameworks such as TensorFlow or PyTorch. The input is numerical data representing operational characteristics, and the output is analysis results related to the suitability evaluation of the robots. Based on these analysis results, the server evaluates the efficiency of each robot.
[0125] Step 5:
[0126] The terminal (administrator's tablet or PC) functions as hardware that displays questions sent from the server and prompts the administrator to answer. The input is question data from the server, and the output is answer data from the administrator. The specific actions performed on the terminal are the presentation of questions and the acceptance of answers.
[0127] Step 6:
[0128] The server uses the response data to the questions to perform an aptitude assessment. The input consists of the analysis results of operational characteristics and the response data, and the output is the assessed aptitude data. Based on this data, the server calculates the optimal role.
[0129] Step 7:
[0130] The server utilizes a generated AI model based on the suitability assessment results to propose the most efficient robot role. The input is suitability data, and the output is information about the proposed role. The server transmits this information to the factory management system.
[0131] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0132] The present invention is a comprehensive sports aptitude evaluation system that utilizes video recognition technology and emotion recognition technology, with the aim of supporting users in performing sports in the optimal position. The embodiments for carrying out the present invention will be described in detail below.
[0133] The user first records their sports play in video format using a recording device. This video is uploaded from the user's device to a server via the network. The server analyzes the received video as described above to extract characteristics of the user's movements. Furthermore, using an emotion engine, it simultaneously recognizes changes in the user's emotions through facial expressions and voice from the video. This allows for the evaluation to incorporate not only movements but also the emotional aspects of play.
[0134] In parallel, users answer questions about sports. This interaction takes place on the device, and detailed answers from the user are sent to the server via a question form. These answers provide information about the user's self-assessment and playing style, and, combined with emotional and behavioral data, form the basis for further analysis.
[0135] The server combines behavioral feature data, emotional state data, and question response data, and uses a machine learning model to evaluate the most suitable position for the user. This evaluation process particularly considers emotional data, analyzing which position elicits the most positive emotions from the user. This enables balanced position recommendations that also take into account the user's psychological condition.
[0136] The evaluation results are sent from the server to the user's terminal, which then visually displays them to the user. For example, the server can detect increases in the user's stress levels from their gameplay video and suggest a more relaxed playing position accordingly, allowing the user to concentrate better and perform more effectively.
[0137] In this way, the system takes into account both the user's physical characteristics and emotional factors, providing strong support for improving and optimizing their playing style in sports.
[0138] The following describes the processing flow.
[0139] Step 1:
[0140] The user records their sports play with a recording device and saves the video data to their device. Next, the user uses the device's interface to upload the recorded video to the system's server. The server saves the received video to its storage.
[0141] Step 2:
[0142] The server processes the saved video frame by frame and extracts motion features using an image recognition algorithm. In this process, the server extracts motion data for each frame (e.g., speed of movement, body angle) as numerical data and temporarily stores it in a database.
[0143] Step 3:
[0144] The server then activates an emotion engine that analyzes the user's facial expressions and voice tone from the video data to recognize the user's emotional state. The recognized emotions are extracted as indicators of stress levels and concentration during gameplay and stored along with the action data.
[0145] Step 4:
[0146] Users answer questions about their sports experience and playing style. The device displays an input form for the questions, where users provide answers through comments and multiple-choice options. The device sends these answers to the server, which adds them to a database.
[0147] Step 5:
[0148] The server integrates behavioral feature data, emotional data, and user response data, and inputs them into a machine learning model to assess the user's suitability. Here, the server uses the model to analyze each dataset and identify the optimal sports position based on the user's physical characteristics and emotional performance.
[0149] Step 6:
[0150] The server generates position suggestions based on the model's execution results and sends them to the user's terminal. The terminal receives this data and presents it to the user in an easy-to-read format. The user can review these suggestions and use them to adjust their own training strategy.
[0151] (Example 2)
[0152] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0153] In sports, position selection for athletes is often based primarily on physical characteristics, with a lack of comprehensive evaluation and suggestions that consider emotional state and psychological condition. This makes it difficult to find the optimal position that allows athletes to perform at their best.
[0154] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0155] In this invention, the server includes signal acquisition means for inputting video signals acquired from a camera that records the user's movements, feature extraction means for extracting characteristics of the user's movements based on the video signals, and emotion recognition means for identifying emotional states from facial expressions and sounds in the video. This makes it possible to suggest the optimal sports position that takes into account both the user's movement characteristics and emotional state.
[0156] "Signal acquisition means" refers to means of inputting video signals acquired from a camera used to record the user's actions.
[0157] A "feature extraction means" is a means for extracting the characteristics of a user's actions based on a video signal.
[0158] "Emotion recognition means" refers to methods for identifying an emotional state from the user's facial expressions and voice in a video.
[0159] A "means for receiving responses" refers to a method for presenting questions to users and receiving their answers.
[0160] "Evaluation methods" refer to means of evaluating a user's suitability by combining characteristic data, sentiment data, and response data.
[0161] The "proposal method" refers to a means of suggesting the optimal sports position to the user using a generative AI model based on the evaluated aptitude.
[0162] This invention is a comprehensive sports aptitude evaluation system that utilizes video recognition technology and emotion recognition technology, with the aim of supporting users in performing sports in the optimal position. This system is configured as follows.
[0163] First, users record their sports play in video format using recording devices such as smartphones or video cameras. The recorded videos are saved on the user's device and then uploaded to a server via the network. The server acquires this video signal and extracts action features using video recognition technology. This process uses OpenCV, an open-source video analysis library. The server also analyzes the user's facial expressions and voice in the video using Google Cloud's sentiment analysis API and other tools to recognize their emotional state.
[0164] Next, the user answers questions about sports on their device, and this answer data is sent to the server. This question-and-answer data provides detailed information about the user's self-assessment and playing style, and, combined with emotion and behavior data, forms the basis for further analysis.
[0165] The server integrates this data and uses a generative AI model to evaluate the most suitable sports position for the user. This evaluation process takes the user's emotional data into particular, making it possible to suggest a balanced position that reflects their psychological condition. For example, the server can detect increases in stress from the user's gameplay video and suggest a position that allows the user to relax accordingly.
[0166] The evaluation results are sent from the server to the user's terminal and presented visually. An example of a prompt message here is the command, "Suggest the optimal position based on the motion and emotion data extracted from the user's video." In this way, the system considers both the user's physical characteristics and emotional elements, providing strong support for improving and optimizing playing style in sports.
[0167] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0168] Step 1:
[0169] Users record their sports play in video format using their smartphones or video cameras. The input is the user's actions during play, and the output is the recorded video data. This process involves using the device's camera function to start recording and capturing specific plays.
[0170] Step 2:
[0171] The user uploads recorded video data from their device to the server via the internet. The input is the video data stored on the user's device, and the output is the video being sent to the server. This process involves pressing the "Upload" button within the application and establishing a connection with the server.
[0172] Step 3:
[0173] The server receives uploaded video data and extracts motion features using video recognition technology. The input is video data, and the output is motion feature data. Specifically, this involves analyzing body movements using the OpenCV library and calculating joint positions and movement patterns.
[0174] Step 4:
[0175] The server uses emotion recognition tools to identify the emotional state of the user from their facial expressions and voice in a video. The input is video data, and the output is emotional state data. Specifically, this involves using an emotion analysis API to analyze changes in facial expressions and voice tone to identify the type of emotion.
[0176] Step 5:
[0177] The user answers a sports-related question form on their device and sends the answers to the server. The input is the user's answers, and the output is the response data transferred to the server. The specific actions include entering information into the form and pressing the "Submit" button.
[0178] Step 6:
[0179] The server integrates and analyzes behavioral characteristic data, emotional state data, and user question response data. The input includes this data, and the output is the integrated analysis results. This includes using a generative AI model to analyze this information and evaluate the optimal position for the user.
[0180] Step 7:
[0181] The server sends the evaluation results to the user's terminal, which then displays the results visually. The input includes the analysis results, and the output is the evaluation results displayed on the terminal for the user to view. This process includes a function to display detailed evaluation information by pressing the "View Results" button.
[0182] (Application Example 2)
[0183] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0184] In conventional factory work, it has been difficult to suggest the optimal work position while considering not only the worker's physical movements but also their emotional changes, resulting in decreased work efficiency and increased psychological burden. There is a need for a system that comprehensively evaluates the physical and mental state of workers in the work environment and recommends the most effective and least stressful work position and tasks.
[0185] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0186] In this invention, the server includes signal acquisition means for inputting video signals acquired from a camera for recording user movements, feature extraction means for extracting characteristics of the user's movements based on the video signals, and emotion recognition means for recognizing the user's emotions and generating emotion data. This makes it possible to analyze the user's movements and emotions in real time and propose the optimal work position and tasks.
[0187] A "signal acquisition means" is a device or module that has the function of inputting video signals acquired from a camera into the system.
[0188] A "feature extraction means" is a device or system that analyzes acquired video signals and has the function of extracting the type of user's movements and physical characteristics.
[0189] "Emotion recognition means" refers to devices or systems that have the function of recognizing the emotional state of a user using voice analysis or facial expression analysis and generating emotional data.
[0190] A "response receiving means" refers to a device or module that has the function of presenting questions to users and receiving their responses.
[0191] "Evaluation means" refers to a device or module that has the function of evaluating the user's suitability by combining feature data obtained by feature extraction means, emotion data generated by emotion recognition means, and response data received by response reception means.
[0192] A "suggestion means" is a device or system that has the function of suggesting the optimal work position or task to the user based on the aptitude evaluated by the evaluation means.
[0193] The system for implementing this invention has the function of analyzing the user's actions and emotional state in real time and suggesting the optimal work position and tasks. The server receives video signals transmitted from the user's terminal through signal acquisition means and extracts characteristics of the actions by analyzing them. Furthermore, using emotion recognition means, it recognizes emotions from the user's face and voice and generates emotion data.
[0194] The terminal presents the user with questions about their work and receives their answers through a response reception mechanism. By combining these answers with data obtained by feature extraction and emotion recognition mechanisms, an evaluation mechanism within the server performs a comprehensive suitability assessment, and the suggestion mechanism uses the evaluation results to propose the optimal work position and tasks. This system enables the provision of a more balanced task assignment and work environment that reflects the user's behavior patterns and emotional state.
[0195] For example, consider a situation in a manufacturing plant where a worker is standing for long periods. This system can detect the user's fatigue and stress in real time and, based on that, suggest the optimal position and relatively lighter tasks for them. This makes it possible to reduce the worker's workload while maintaining productivity.
[0196] An example of a prompt message is, "Based on the worker's current emotional state, which task or position is appropriate?" Based on this prompt message, the AI model evaluates and generates suggestions.
[0197] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0198] Step 1:
[0199] The server receives video signals transmitted from the user's terminal using signal acquisition means. The video signal is provided as input, and preprocessing is performed based on it for motion analysis. This preprocessing removes noise from the video and converts it into an analyzable format.
[0200] Step 2:
[0201] Using the processed video data, the server's feature extraction mechanism extracts motion characteristics. Pre-processed video data is used as input, and the user's motion patterns are generated as output. Specifically, the type and frequency of movement, the trajectory of body movements, etc., are analyzed by an algorithm.
[0202] Step 3:
[0203] Simultaneously, the server's emotion recognition system analyzes the user's face and voice in the video to generate emotion data. Voice and facial feature data are used as input. Based on facial expressions and tone of voice, it recognizes emotions such as satisfaction, dissatisfaction, and stress.
[0204] Step 4:
[0205] The terminal presents the user with questions about specific tasks and receives the user's answers through a response reception mechanism. The input includes the answers selected by the user on the terminal and is transmitted directly to the server. The questions often concern the user's experience and level of comfort.
[0206] Step 5:
[0207] The server integrates data obtained by feature extraction and sentiment recognition methods, along with user response data, using evaluation methods to assess the user's suitability. All collected data is included as input, and the output generates recommendations for the most suitable work position and tasks for the user.
[0208] Step 6:
[0209] Based on the evaluation results, the server's suggestion mechanism uses a generated AI model to create prompt messages that suggest the optimal work position and tasks for the user. The output includes specific work location suggestions and instructions for the user, and is presented to the user via the terminal.
[0210] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0211] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0212] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0213] [Second Embodiment]
[0214] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0215] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0216] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0217] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0218] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0219] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0220] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0221] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0222] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0223] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0224] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0225] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0226] This system aims to support users in performing sports from the optimal position, and is a sports aptitude evaluation system that utilizes video recognition technology and machine learning. The embodiments for carrying out the present invention will be described in detail below.
[0227] The user first records their sports play in video format using a recording device. This recorded video is then uploaded from the terminal to a server via the network. The server processes the received video and performs motion analysis. Specifically, the server divides the video into frames and uses image recognition technology to extract the motion of each frame as numerical data. This feature data is useful for detailed analysis of the user's movements, physique, form, etc.
[0228] Simultaneously, users are required to answer questions via their device regarding their sports experience, playing style, and preferred movements. The device then sends these answers to a server. This question-and-answer data integrates the user's subjective information with objective data.
[0229] The server combines feature data and question-answer data, and uses a machine learning model to evaluate the most suitable position for the user based on this information. It compares this information with data on athletes obtained from past databases to assess suitability. This makes it possible to recommend the position best suited to the user's movement style and characteristics.
[0230] The recommended results are sent from the server to the user's device, which then displays them to the user. The user can review the suggested results and use them to practice and train for the newly evaluated position. For example, if the server recommends outfielder as the optimal position based on the user's pitching motion and running ability, the user can use that feedback to refine their suitability as an outfielder.
[0231] In this way, this system maximizes the user's potential and supports the improvement and optimization of their playing style in sports.
[0232] The following describes the processing flow.
[0233] Step 1:
[0234] The user records their sports play with a recording device and saves the video file to their device. Using the device's interface, they upload the recorded video to the server. The server receives the uploaded video file and saves it to its storage.
[0235] Step 2:
[0236] The server divides the saved video into frames and applies an image recognition algorithm to each frame to extract motion characteristics. During this process, the server collects information such as the user's posture, movement speed, and angle as numerical data. The extracted data is temporarily stored.
[0237] Step 3:
[0238] The user is asked to answer questions about their sports experience. The device displays a question form, and the user enters information about their experience, playing style, and preferred movements. The device then sends these answers to the server.
[0239] Step 4:
[0240] The server receives image recognition feature data extracted from the frames and question data answered by the user, and integrates this data. The server then performs preprocessing to prepare the data for use in machine learning models.
[0241] Step 5:
[0242] The server inputs the prepared data into a machine learning model and begins the process of evaluating the user's suitability. The model analyzes influential features and performs calculations to recommend the optimal position by comparing them with past data.
[0243] Step 6:
[0244] The server sends the user the optimal sports position derived from machine learning. The user's device receives this information and displays the results in a visually appealing format. Based on this information, the user adjusts their training and plans how to adapt to the suggested position.
[0245] (Example 1)
[0246] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0247] In today's sports environment, quickly and accurately determining the optimal playing position based on a player's characteristics is a challenging task. Traditional methods often rely on subjective evaluations and lack objective, data-driven analysis, which can prevent players from fully realizing their potential.
[0248] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0249] In this invention, the server includes data acquisition means, feature analysis means, and information gathering means. This makes it possible to integrate and evaluate the user's movement characteristics and subjective information, and objectively recommend the optimal movement position.
[0250] "Data acquisition means" refers to a function that inputs video data acquired from a camera that records the user's actions.
[0251] A "feature analysis method" is a function that divides video data into frames and extracts the actions of each frame as numerical data.
[0252] An "information gathering tool" is a function that presents questions to the user and receives the answers in digital format.
[0253] The "analysis means" is a function that integrates numerical data extracted by the feature analysis means with response data received by the information collection means to evaluate the user's characteristics.
[0254] "Recommendation method" refers to a function that recommends the optimal exercise position to the user based on the evaluation results from the analysis method.
[0255] This invention is a system that analyzes the user's movement characteristics based on data and proposes the optimal sports position. Specific embodiments of this system are shown below.
[0256] System configuration:
[0257] The system primarily consists of a server, terminals, and user operations. Users record their sports play in video format using recording devices (e.g., smartphones or dedicated cameras). Users upload these recorded videos to the server via their terminals. The server divides the received video into frames and uses image recognition technology as a feature analysis tool to extract and analyze the actions in each frame as numerical data. In particular, the use of deep learning libraries and computer vision models is conceivable.
[0258] Data collection:
[0259] Users answer questions about their sports experience and preferred playing style through a terminal. This information is entered into the terminal in digital format and then transmitted to a server. The server aggregates a large amount of question-and-answer data through this information collection method.
[0260] Data integration and analysis:
[0261] The server integrates numerical data extracted by feature analysis tools with question response data collected by information gathering tools. Based on the resulting dataset, machine learning techniques are used as an analytical tool to objectively evaluate the user's characteristics. The generative AI model compares the user with a database of past athletes to recommend the most suitable sports position for the user.
[0262] Presentation of results:
[0263] Based on the evaluation results, the server utilizes recommended methods to present the optimal sports position to the user on the device. The device can then display feedback to the user and provide specific training guidelines.
[0264] Specific example:
[0265] For example, if the server analyzes a user's video data and recommends that "outfielder" is the best position based on their pitching motion and excellent running ability, the user can then train based on this evaluation.
[0266] Example prompts for generative AI models:
[0267] "Evaluate the user's pitching and running abilities based on their video data, and recommend a suitable sports position."
[0268] "Consider your experience and playing style, and select the position that best suits you."
[0269] This invention provides a precise approach to maximizing sports performance.
[0270] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0271] Step 1:
[0272] The user records their sports play in video format using a recording device. The input is the user's actions, and video data is generated as output. Specifically, the user uses a smartphone to film their soccer dribbling movements.
[0273] Step 2:
[0274] The device uploads video data to the server. The input is video data, and the output is data transfer to the server via the network. For example, pressing the "upload" button in the dedicated application sends the video to the server.
[0275] Step 3:
[0276] The server divides the received video into frames. The input is the video data transferred to the server, and the output is a dataset of image frames. Specifically, the server divides the video into 30 frames per second and generates a still image for each frame.
[0277] Step 4:
[0278] The server analyzes each frame of the video and uses image recognition technology as a feature analysis tool to extract motion characteristics. The input is image frame data, and the output is numerical data representing the motion of each frame. Specifically, the server detects things like the position of the feet and the angle of the body, and quantifies these.
[0279] Step 5:
[0280] Users answer questions about their sports experience and preferred playing style through their device. The input is user-provided data, and structured response information is generated as output. Specifically, users answer questions in a survey format such as, "What are your preferred movements?"
[0281] Step 6:
[0282] The terminal sends the user's response data to the server. The input is the user's response information, and the output is the transfer to the server. As a specific operation, the user presses the "Send" button on the terminal, and the data is sent to the server.
[0283] Step 7:
[0284] The server integrates the numerical data and response data extracted by the feature analysis means. The input is the digitized operation data and response data, and the output is the integrated data set. Here, comprehensive analysis becomes possible through data integration.
[0285] Step 8:
[0286] The server analyzes using the integrated data by the generated AI model to evaluate the characteristics of the user. The input is the integrated data set, and the output is the proposal as the evaluation result. As a specific example, the server analyzes the operation pattern with a machine learning algorithm and recommends the optimal position.
[0287] Step 9:
[0288] The server sends the result to the terminal, and the terminal presents the optimal position to the user. The input is the evaluation result sent from the server, and the output is the feedback display to the user. The user can perform training based on this feedback.
[0289] (Application Example 1)
[0290] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0291] There is a demand to improve the efficiency of robots that perform multiple tasks in a factory. However, currently, it is difficult to grasp the optimal working position and role of each robot, and the issue is to arrange them in the right place for the right job. This leads to a decrease in productivity and a problem of wasting resources.
[0292] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0293] In this invention, the server includes data acquisition means for inputting video signals obtained from a camera for recording user movements, data analysis means for extracting characteristics of the user's movements based on the video signals, and information receiving means for presenting questions to the user and receiving their answers. This makes it possible to recommend the optimal role and position for each robot to perform its work most efficiently.
[0294] The "data acquisition means" is responsible for acquiring video signals from devices that record the movements of users or robots, and transmitting them to a server for analysis.
[0295] A "data analysis means" is a device that uses a visual data recognition algorithm to extract motion characteristics from acquired video signals and analyzes them as numerical data.
[0296] An "information reception method" is a system that presents questions to users or administrators, receives their answers, and transmits them to a server as data for use in suitability assessments.
[0297] "Suitability evaluation means" refers to a system that integrates feature data obtained from data analysis means and information data acquired from information reception means to evaluate the most suitable work position and role for each robot.
[0298] The "role proposal method" is a process that proposes the most suitable work role for a robot based on the evaluation results derived from the aptitude evaluation method, and supports improvements in placement and work content.
[0299] This system is designed to optimize the work efficiency of robots in factories. First, cameras are attached to the robots in the factory, and these cameras continuously record the robots' movements. This serves as a means of data acquisition.
[0300] The server uses video recognition technologies such as OpenCV to divide the acquired video signal into frames and processes the motion of each frame using data analysis tools. This process extracts the robot's motion characteristics and external features as numerical data. Furthermore, the server uses machine learning frameworks such as TensorFlow and PyTorch to analyze the data.
[0301] As a means of receiving information, administrators are provided with tablets or PCs, on which questions from the server are presented. Administrators answer these questions, and the content of their answers is sent to the server. The server uses these answers to evaluate the robot's suitability using an aptitude evaluation system.
[0302] As a result, the optimal robot work position and role are suggested by the role suggestion system. This information can be viewed on the factory's management system.
[0303] For example, if a suitability assessment recommends that a particular robot be specialized for transport tasks in order to improve the efficiency of assembly and transport operations, the factory manager can change the robot's placement based on that suggestion.
[0304] An example of a prompt in a generative AI model is, "Based on the operation data of the factory robot, please suggest the most efficient work role."
[0305] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0306] Step 1:
[0307] A factory robot equipped with a camera records its own operations during work as videos. The input is real-time video data. This video data is later uploaded to a server as data acquisition means.
[0308] Step 2:
[0309] The server divides the uploaded video data frame by frame. The input is video data, and the output is a collection of still images divided into each frame. The server uses this divided frame to prepare for extracting motion characteristics using OpenCV.
[0310] Step 3:
[0311] The server, as data analysis means, extracts the motion characteristics of the robot from each frame as numerical data using OpenCV. The input is the divided frame, and the output is numerical data indicating the motion characteristics. The server uses this numerical data for subsequent analysis.
[0312] Step 4:
[0313] The server analyzes the numerical data using a machine learning framework such as TensorFlow or PyTorch. The input is numerical data indicating the motion characteristics, and the output is an analysis result related to the suitability evaluation of the robot. The server evaluates the efficiency of each robot based on this analysis result.
[0314] Step 5:
[0315] [[ID=三十二]]The terminal (the administrator's tablet or PC) functions as hardware that displays the questions sent from the server and prompts the administrator to answer. The input is question data from the server, and the output is answer data from the administrator. The specific operations for the terminal are presenting questions and receiving answers.
[0316] Step 6:
[0317] The server uses the response data to the questions to perform an aptitude assessment. The input consists of the analysis results of operational characteristics and the response data, and the output is the assessed aptitude data. Based on this data, the server calculates the optimal role.
[0318] Step 7:
[0319] The server utilizes a generated AI model based on the suitability assessment results to propose the most efficient robot role. The input is suitability data, and the output is information about the proposed role. The server transmits this information to the factory management system.
[0320] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0321] The present invention is a comprehensive sports aptitude evaluation system that utilizes video recognition technology and emotion recognition technology, with the aim of supporting users in performing sports in the optimal position. The embodiments for carrying out the present invention will be described in detail below.
[0322] The user first records their sports play in video format using a recording device. This video is uploaded from the user's device to a server via the network. The server analyzes the received video as described above to extract characteristics of the user's movements. Furthermore, using an emotion engine, it simultaneously recognizes changes in the user's emotions through facial expressions and voice from the video. This allows for the evaluation to incorporate not only movements but also the emotional aspects of play.
[0323] In parallel, users answer questions about sports. This interaction takes place on the device, and detailed answers from the user are sent to the server via a question form. These answers provide information about the user's self-assessment and playing style, and, combined with emotional and behavioral data, form the basis for further analysis.
[0324] The server combines behavioral feature data, emotional state data, and question response data, and uses a machine learning model to evaluate the most suitable position for the user. This evaluation process particularly considers emotional data, analyzing which position elicits the most positive emotions from the user. This enables balanced position recommendations that also take into account the user's psychological condition.
[0325] The evaluation results are sent from the server to the user's terminal, which then visually displays them to the user. For example, the server can detect increases in the user's stress levels from their gameplay video and suggest a more relaxed playing position accordingly, allowing the user to concentrate better and perform more effectively.
[0326] In this way, the system takes into account both the user's physical characteristics and emotional factors, providing strong support for improving and optimizing their playing style in sports.
[0327] The following describes the processing flow.
[0328] Step 1:
[0329] The user records their sports play with a recording device and saves the video data to their device. Next, the user uses the device's interface to upload the recorded video to the system's server. The server saves the received video to its storage.
[0330] Step 2:
[0331] The server processes the saved video frame by frame and extracts motion features using an image recognition algorithm. In this process, the server extracts motion data for each frame (e.g., speed of movement, body angle) as numerical data and temporarily stores it in a database.
[0332] Step 3:
[0333] The server then activates an emotion engine that analyzes the user's facial expressions and voice tone from the video data to recognize the user's emotional state. The recognized emotions are extracted as indicators of stress levels and concentration during gameplay and stored along with the action data.
[0334] Step 4:
[0335] Users answer questions about their sports experience and playing style. The device displays an input form for the questions, where users provide answers through comments and multiple-choice options. The device sends these answers to the server, which adds them to a database.
[0336] Step 5:
[0337] The server integrates behavioral feature data, emotional data, and user response data, and inputs them into a machine learning model to assess the user's suitability. Here, the server uses the model to analyze each dataset and identify the optimal sports position based on the user's physical characteristics and emotional performance.
[0338] Step 6:
[0339] The server generates position suggestions based on the model's execution results and sends them to the user's terminal. The terminal receives this data and presents it to the user in an easy-to-read format. The user can review these suggestions and use them to adjust their own training strategy.
[0340] (Example 2)
[0341] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0342] In sports, position selection for athletes is often based primarily on physical characteristics, with a lack of comprehensive evaluation and suggestions that consider emotional state and psychological condition. This makes it difficult to find the optimal position that allows athletes to perform at their best.
[0343] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0344] In this invention, the server includes signal acquisition means for inputting video signals acquired from a camera that records the user's movements, feature extraction means for extracting characteristics of the user's movements based on the video signals, and emotion recognition means for identifying emotional states from facial expressions and sounds in the video. This makes it possible to suggest the optimal sports position that takes into account both the user's movement characteristics and emotional state.
[0345] "Signal acquisition means" refers to means of inputting video signals acquired from a camera used to record the user's actions.
[0346] A "feature extraction means" is a means for extracting the characteristics of a user's actions based on a video signal.
[0347] "Emotion recognition means" refers to methods for identifying an emotional state from the user's facial expressions and voice in a video.
[0348] A "means for receiving responses" refers to a method for presenting questions to users and receiving their answers.
[0349] "Evaluation methods" refer to means of evaluating a user's suitability by combining characteristic data, sentiment data, and response data.
[0350] The "proposal method" refers to a means of suggesting the optimal sports position to the user using a generative AI model based on the evaluated aptitude.
[0351] This invention is a comprehensive sports aptitude evaluation system that utilizes video recognition technology and emotion recognition technology, with the aim of supporting users in performing sports in the optimal position. This system is configured as follows.
[0352] First, users record their sports performance in video format using recording devices such as smartphones or video cameras. The recorded videos are saved on the user's device and then uploaded to a server via the network. The server acquires this video signal and extracts action features using video recognition technology. This process uses OpenCV, an open-source video analysis library. The server also analyzes the user's facial expressions and voice in the video using Google Cloud's sentiment analysis API and other tools to recognize their emotional state.
[0353] Next, the user answers questions about sports on their device, and this answer data is sent to the server. This question-and-answer data provides detailed information about the user's self-assessment and playing style, and, combined with emotion and behavior data, forms the basis for further analysis.
[0354] The server integrates this data and uses a generative AI model to evaluate the most suitable sports position for the user. This evaluation process takes the user's emotional data into particular, making it possible to suggest a balanced position that reflects their psychological condition. For example, the server can detect increases in stress from the user's gameplay video and suggest a position that allows the user to relax accordingly.
[0355] The evaluation results are sent from the server to the user's terminal and presented visually. An example of a prompt message here is the command, "Suggest the optimal position based on the motion and emotion data extracted from the user's video." In this way, the system considers both the user's physical characteristics and emotional elements, providing strong support for improving and optimizing playing style in sports.
[0356] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0357] Step 1:
[0358] Users record their sports play in video format using their smartphones or video cameras. The input is the user's actions during play, and the output is the recorded video data. This process involves using the device's camera function to start recording and capturing specific plays.
[0359] Step 2:
[0360] The user uploads recorded video data from their device to the server via the internet. The input is the video data stored on the user's device, and the output is the video being sent to the server. This process involves pressing the "Upload" button within the application and establishing a connection with the server.
[0361] Step 3:
[0362] The server receives uploaded video data and extracts motion features using video recognition technology. The input is video data, and the output is motion feature data. Specifically, this involves analyzing body movements using the OpenCV library and calculating joint positions and movement patterns.
[0363] Step 4:
[0364] The server uses emotion recognition tools to identify the emotional state of the user from their facial expressions and voice in a video. The input is video data, and the output is emotional state data. Specifically, this involves using an emotion analysis API to analyze changes in facial expressions and voice tone to identify the type of emotion.
[0365] Step 5:
[0366] The user answers a sports-related question form on their device and sends the answers to the server. The input is the user's answers, and the output is the response data transferred to the server. The specific actions include entering information into the form and pressing the "Submit" button.
[0367] Step 6:
[0368] The server integrates and analyzes behavioral characteristic data, emotional state data, and user question response data. The input includes this data, and the output is the integrated analysis results. This includes using a generative AI model to analyze this information and evaluate the optimal position for the user.
[0369] Step 7:
[0370] The server sends the evaluation results to the user's terminal, which then displays the results visually. The input includes the analysis results, and the output is the evaluation results displayed on the terminal for the user to view. This process includes a function to display detailed evaluation information by pressing the "View Results" button.
[0371] (Application Example 2)
[0372] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0373] In conventional factory work, it has been difficult to suggest the optimal work position while considering not only the worker's physical movements but also their emotional changes, resulting in decreased work efficiency and increased psychological burden. There is a need for a system that comprehensively evaluates the physical and mental state of workers in the work environment and recommends the most effective and least stressful work position and tasks.
[0374] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0375] In this invention, the server includes signal acquisition means for inputting video signals acquired from a camera for recording user movements, feature extraction means for extracting characteristics of the user's movements based on the video signals, and emotion recognition means for recognizing the user's emotions and generating emotion data. This makes it possible to analyze the user's movements and emotions in real time and propose the optimal work position and tasks.
[0376] A "signal acquisition means" is a device or module that has the function of inputting video signals acquired from a camera into the system.
[0377] A "feature extraction means" is a device or system that analyzes acquired video signals and has the function of extracting the type of user's movements and physical characteristics.
[0378] "Emotion recognition means" refers to devices or systems that have the function of recognizing the emotional state of a user using voice analysis or facial expression analysis and generating emotional data.
[0379] A "response receiving means" refers to a device or module that has the function of presenting questions to users and receiving their responses.
[0380] "Evaluation means" refers to a device or module that has the function of evaluating the user's suitability by combining feature data obtained by feature extraction means, emotion data generated by emotion recognition means, and response data received by response reception means.
[0381] A "suggestion means" is a device or system that has the function of suggesting the optimal work position or task to the user based on the aptitude evaluated by the evaluation means.
[0382] The system for implementing this invention has the function of analyzing the user's actions and emotional state in real time and suggesting the optimal work position and tasks. The server receives video signals transmitted from the user's terminal through signal acquisition means and extracts characteristics of the actions by analyzing them. Furthermore, using emotion recognition means, it recognizes emotions from the user's face and voice and generates emotion data.
[0383] The terminal presents the user with questions about their work and receives their answers through a response reception mechanism. By combining these answers with data obtained by feature extraction and emotion recognition mechanisms, an evaluation mechanism within the server performs a comprehensive suitability assessment, and the suggestion mechanism uses the evaluation results to propose the optimal work position and tasks. This system enables the provision of a more balanced task assignment and work environment that reflects the user's behavior patterns and emotional state.
[0384] For example, consider a situation in a manufacturing plant where a worker is standing for long periods. This system can detect the user's fatigue and stress in real time and, based on that, suggest the optimal position and relatively lighter tasks for them. This makes it possible to reduce the worker's workload while maintaining productivity.
[0385] An example of a prompt message is, "Based on the worker's current emotional state, which task or position is appropriate?" Based on this prompt message, the AI model evaluates and generates suggestions.
[0386] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0387] Step 1:
[0388] The server receives video signals transmitted from the user's terminal using signal acquisition means. The video signal is provided as input, and preprocessing is performed based on it for motion analysis. This preprocessing removes noise from the video and converts it into an analyzable format.
[0389] Step 2:
[0390] Using the processed video data, the server's feature extraction mechanism extracts motion characteristics. Pre-processed video data is used as input, and the user's motion patterns are generated as output. Specifically, the type and frequency of movement, the trajectory of body movements, etc., are analyzed by an algorithm.
[0391] Step 3:
[0392] Simultaneously, the server's emotion recognition system analyzes the user's face and voice in the video to generate emotion data. Voice and facial feature data are used as input. Based on facial expressions and tone of voice, it recognizes emotions such as satisfaction, dissatisfaction, and stress.
[0393] Step 4:
[0394] The terminal presents the user with questions about specific tasks and receives the user's answers through a response reception mechanism. The input includes the answers selected by the user on the terminal and is transmitted directly to the server. The questions often concern the user's experience and level of comfort.
[0395] Step 5:
[0396] The server integrates data obtained by feature extraction and sentiment recognition methods, along with user response data, using evaluation methods to assess the user's suitability. All collected data is included as input, and the output generates recommendations for the most suitable work position and tasks for the user.
[0397] Step 6:
[0398] Based on the evaluation results, the server's suggestion mechanism uses a generated AI model to create prompt messages that suggest the optimal work position and tasks for the user. The output includes specific work location suggestions and instructions for the user, and is presented to the user via the terminal.
[0399] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0400] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0401] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0402] [Third Embodiment]
[0403] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0404] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0405] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0406] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0407] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0408] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0409] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0410] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0411] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0412] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0413] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0414] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0415] This system aims to support users in performing sports from the optimal position, and is a sports aptitude evaluation system that utilizes video recognition technology and machine learning. The embodiments for carrying out the present invention will be described in detail below.
[0416] The user first records their sports play in video format using a recording device. This recorded video is then uploaded from the terminal to a server via the network. The server processes the received video and performs motion analysis. Specifically, the server divides the video into frames and uses image recognition technology to extract the motion of each frame as numerical data. This feature data is useful for detailed analysis of the user's movements, physique, form, etc.
[0417] Simultaneously, users are required to answer questions via their device regarding their sports experience, playing style, and preferred movements. The device then sends these answers to a server. This question-and-answer data integrates the user's subjective information with objective data.
[0418] The server combines feature data and question-answer data, and uses a machine learning model to evaluate the most suitable position for the user based on this information. It compares this information with data on athletes obtained from past databases to assess suitability. This makes it possible to recommend the position best suited to the user's movement style and characteristics.
[0419] The recommended results are sent from the server to the user's device, which then displays them to the user. The user can review the suggested results and use them to practice and train for the newly evaluated position. For example, if the server recommends outfielder as the optimal position based on the user's pitching motion and running ability, the user can use that feedback to refine their suitability as an outfielder.
[0420] In this way, this system maximizes the user's potential and supports the improvement and optimization of their playing style in sports.
[0421] The following describes the processing flow.
[0422] Step 1:
[0423] The user records their sports play with a recording device and saves the video file to their device. Using the device's interface, they upload the recorded video to the server. The server receives the uploaded video file and saves it to its storage.
[0424] Step 2:
[0425] The server divides the saved video into frames and applies an image recognition algorithm to each frame to extract motion characteristics. During this process, the server collects information such as the user's posture, movement speed, and angle as numerical data. The extracted data is temporarily stored.
[0426] Step 3:
[0427] The user is asked to answer questions about their sports experience. The device displays a question form, and the user enters information about their experience, playing style, and preferred movements. The device then sends these answers to the server.
[0428] Step 4:
[0429] The server receives image recognition feature data extracted from the frames and question data answered by the user, and integrates this data. The server then performs preprocessing to prepare the data for use in machine learning models.
[0430] Step 5:
[0431] The server inputs the prepared data into a machine learning model and begins the process of evaluating the user's suitability. The model analyzes influential features and performs calculations to recommend the optimal position by comparing them with past data.
[0432] Step 6:
[0433] The server sends the user the optimal sports position derived from machine learning. The user's device receives this information and displays the results in a visually appealing format. Based on this information, the user adjusts their training and plans how to adapt to the suggested position.
[0434] (Example 1)
[0435] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0436] In today's sports environment, quickly and accurately determining the optimal playing position based on a player's characteristics is a challenging task. Traditional methods often rely on subjective evaluations and lack objective, data-driven analysis, which can prevent players from fully realizing their potential.
[0437] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0438] In this invention, the server includes data acquisition means, feature analysis means, and information gathering means. This makes it possible to integrate and evaluate the user's movement characteristics and subjective information, and objectively recommend the optimal movement position.
[0439] "Data acquisition means" refers to a function that inputs video data acquired from a camera that records the user's actions.
[0440] A "feature analysis method" is a function that divides video data into frames and extracts the actions of each frame as numerical data.
[0441] An "information gathering tool" is a function that presents questions to the user and receives the answers in digital format.
[0442] The "analysis means" is a function that integrates numerical data extracted by the feature analysis means with response data received by the information collection means to evaluate the user's characteristics.
[0443] "Recommendation method" refers to a function that recommends the optimal exercise position to the user based on the evaluation results from the analysis method.
[0444] This invention is a system that analyzes the user's movement characteristics based on data and proposes the optimal sports position. Specific embodiments of this system are shown below.
[0445] System configuration:
[0446] The system primarily consists of a server, terminals, and user operations. Users record their sports play in video format using recording devices (e.g., smartphones or dedicated cameras). Users upload these recorded videos to the server via their terminals. The server divides the received video into frames and uses image recognition technology as a feature analysis tool to extract and analyze the actions in each frame as numerical data. In particular, the use of deep learning libraries and computer vision models is conceivable.
[0447] Data collection:
[0448] Users answer questions about their sports experience and preferred playing style through a terminal. This information is entered into the terminal in digital format and then transmitted to a server. The server aggregates a large amount of question-and-answer data through this information collection method.
[0449] Data integration and analysis:
[0450] The server integrates numerical data extracted by feature analysis tools with question response data collected by information gathering tools. Based on the resulting dataset, machine learning techniques are used as an analytical tool to objectively evaluate the user's characteristics. The generative AI model compares the user with a database of past athletes to recommend the most suitable sports position for the user.
[0451] Presentation of results:
[0452] Based on the evaluation results, the server utilizes recommended methods to present the optimal sports position to the user on the device. The device can then display feedback to the user and provide specific training guidelines.
[0453] Specific example:
[0454] For example, if the server analyzes a user's video data and recommends that "outfielder" is the best position based on their pitching motion and excellent running ability, the user can then train based on this evaluation.
[0455] Example prompts for generative AI models:
[0456] "Evaluate the user's pitching and running abilities based on their video data, and recommend a suitable sports position."
[0457] "Consider your experience and playing style, and select the position that best suits you."
[0458] This invention provides a precise approach to maximizing sports performance.
[0459] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0460] Step 1:
[0461] The user records their sports play in video format using a recording device. The input is the user's actions, and video data is generated as output. Specifically, the user uses a smartphone to film their soccer dribbling movements.
[0462] Step 2:
[0463] The device uploads video data to the server. The input is video data, and the output is data transfer to the server via the network. For example, pressing the "upload" button in the dedicated application sends the video to the server.
[0464] Step 3:
[0465] The server divides the received video into frames. The input is the video data transferred to the server, and the output is a dataset of image frames. Specifically, the server divides the video into 30 frames per second and generates a still image for each frame.
[0466] Step 4:
[0467] The server analyzes each frame of the video and uses image recognition technology as a feature analysis tool to extract motion characteristics. The input is image frame data, and the output is numerical data representing the motion of each frame. Specifically, the server detects things like the position of the feet and the angle of the body, and quantifies these.
[0468] Step 5:
[0469] Users answer questions about their sports experience and preferred playing style through their device. The input is user-provided data, and structured response information is generated as output. Specifically, users answer questions in a survey format such as, "What are your preferred movements?"
[0470] Step 6:
[0471] The terminal sends the user's response data to the server. The input is the user's response information, and the output is the transfer to the server. Specifically, the user presses the "Send" button on the terminal, and the data is sent to the server.
[0472] Step 7:
[0473] The server integrates numerical data extracted by the feature analysis method with response data. The input consists of quantified behavioral data and response data, and the output is an integrated dataset. This data integration enables comprehensive analysis.
[0474] Step 8:
[0475] The server uses integrated data to perform analysis with a generative AI model and evaluate the user's characteristics. The input is the integrated dataset, and the output is a proposal as an evaluation result. Specifically, the server analyzes behavioral patterns using a machine learning algorithm and recommends the optimal position.
[0476] Step 9:
[0477] The server sends the results to the terminal, which then displays the optimal position to the user. The input is the evaluation result sent from the server, and the output is the feedback displayed to the user. The user can then use this feedback to train.
[0478] (Application Example 1)
[0479] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0480] There is a need to improve the efficiency of robots that perform multiple tasks in factories. However, currently, it is difficult to understand the optimal work position and role for each robot, making it challenging to place them in the right place. This leads to decreased productivity and wasted resources.
[0481] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0482] In this invention, the server includes data acquisition means for inputting video signals obtained from a camera for recording user movements, data analysis means for extracting characteristics of the user's movements based on the video signals, and information receiving means for presenting questions to the user and receiving their answers. This makes it possible to recommend the optimal role and position for each robot to perform its work most efficiently.
[0483] The "data acquisition means" is responsible for acquiring video signals from devices that record the movements of users or robots, and transmitting them to a server for analysis.
[0484] A "data analysis means" is a device that uses a visual data recognition algorithm to extract motion characteristics from acquired video signals and analyzes them as numerical data.
[0485] An "information reception method" is a system that presents questions to users or administrators, receives their answers, and transmits them to a server as data for use in suitability assessments.
[0486] "Suitability evaluation means" refers to a system that integrates feature data obtained from data analysis means and information data acquired from information reception means to evaluate the most suitable work position and role for each robot.
[0487] The "role proposal method" is a process that proposes the most suitable work role for a robot based on the evaluation results derived from the aptitude evaluation method, and supports improvements in placement and work content.
[0488] This system is designed to optimize the work efficiency of robots in factories. First, cameras are attached to the robots in the factory, and these cameras continuously record the robots' movements. This serves as a means of data acquisition.
[0489] The server uses video recognition technologies such as OpenCV to divide the acquired video signal into frames and processes the motion of each frame using data analysis tools. This process extracts the robot's motion characteristics and external features as numerical data. Furthermore, the server uses machine learning frameworks such as TensorFlow and PyTorch to analyze the data.
[0490] As a means of receiving information, administrators are provided with tablets or PCs, on which questions from the server are presented. Administrators answer these questions, and the content of their answers is sent to the server. The server uses these answers to evaluate the robot's suitability using an aptitude evaluation system.
[0491] As a result, the optimal robot work position and role are suggested by the role suggestion system. This information can be viewed on the factory's management system.
[0492] For example, if a suitability assessment recommends that a particular robot be specialized for transport tasks in order to improve the efficiency of assembly and transport operations, the factory manager can change the robot's placement based on that suggestion.
[0493] An example of a prompt in a generative AI model is, "Based on the operation data of the factory robot, please suggest the most efficient work role."
[0494] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0495] Step 1:
[0496] A factory robot equipped with a camera records its own movements as video during work. The input is real-time video data. This video data is later uploaded to a server as a data acquisition method.
[0497] Step 2:
[0498] The server divides the uploaded video data into frames. The input is the video data, and the output is a collection of still images divided into each frame. The server uses these divided frames to prepare for extracting the operating characteristics using OpenCV.
[0499] Step 3:
[0500] The server uses OpenCV as a data analysis tool to extract numerical data representing the robot's motion characteristics from each frame. The input is the divided frames, and the output is numerical data showing the motion characteristics. The server then uses this numerical data for subsequent analysis.
[0501] Step 4:
[0502] The server analyzes numerical data using machine learning frameworks such as TensorFlow or PyTorch. The input is numerical data representing operational characteristics, and the output is analysis results related to the suitability evaluation of the robots. Based on these analysis results, the server evaluates the efficiency of each robot.
[0503] Step 5:
[0504] The terminal (administrator's tablet or PC) functions as hardware that displays questions sent from the server and prompts the administrator to answer. The input is question data from the server, and the output is answer data from the administrator. The specific actions performed on the terminal are the presentation of questions and the acceptance of answers.
[0505] Step 6:
[0506] The server uses the response data to the questions to perform an aptitude assessment. The input consists of the analysis results of operational characteristics and the response data, and the output is the assessed aptitude data. Based on this data, the server calculates the optimal role.
[0507] Step 7:
[0508] The server utilizes a generated AI model based on the suitability assessment results to propose the most efficient robot role. The input is suitability data, and the output is information about the proposed role. The server transmits this information to the factory management system.
[0509] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0510] The present invention is a comprehensive sports aptitude evaluation system that utilizes video recognition technology and emotion recognition technology, with the aim of supporting users in performing sports in the optimal position. The embodiments for carrying out the present invention will be described in detail below.
[0511] The user first records their sports play in video format using a recording device. This video is uploaded from the user's device to a server via the network. The server analyzes the received video as described above to extract characteristics of the user's movements. Furthermore, using an emotion engine, it simultaneously recognizes changes in the user's emotions through facial expressions and voice from the video. This allows for the evaluation to incorporate not only movements but also the emotional aspects of play.
[0512] In parallel, users answer questions about sports. This interaction takes place on the device, and detailed answers from the user are sent to the server via a question form. These answers provide information about the user's self-assessment and playing style, and, combined with emotional and behavioral data, form the basis for further analysis.
[0513] The server combines behavioral feature data, emotional state data, and question response data, and uses a machine learning model to evaluate the most suitable position for the user. This evaluation process particularly considers emotional data, analyzing which position elicits the most positive emotions from the user. This enables balanced position recommendations that also take into account the user's psychological condition.
[0514] The evaluation results are sent from the server to the user's terminal, which then visually displays them to the user. For example, the server can detect increases in the user's stress levels from their gameplay video and suggest a more relaxed playing position accordingly, allowing the user to concentrate better and perform more effectively.
[0515] In this way, the system takes into account both the user's physical characteristics and emotional factors, providing strong support for improving and optimizing their playing style in sports.
[0516] The following describes the processing flow.
[0517] Step 1:
[0518] The user records their sports play with a recording device and saves the video data to their device. Next, the user uses the device's interface to upload the recorded video to the system's server. The server saves the received video to its storage.
[0519] Step 2:
[0520] The server processes the saved video frame by frame and extracts motion features using an image recognition algorithm. In this process, the server extracts motion data for each frame (e.g., speed of movement, body angle) as numerical data and temporarily stores it in a database.
[0521] Step 3:
[0522] The server then activates an emotion engine that analyzes the user's facial expressions and voice tone from the video data to recognize the user's emotional state. The recognized emotions are extracted as indicators of stress levels and concentration during gameplay and stored along with the action data.
[0523] Step 4:
[0524] Users answer questions about their sports experience and playing style. The device displays an input form for the questions, where users provide answers through comments and multiple-choice options. The device sends these answers to the server, which adds them to a database.
[0525] Step 5:
[0526] The server integrates behavioral feature data, emotional data, and user response data, and inputs them into a machine learning model to assess the user's suitability. Here, the server uses the model to analyze each dataset and identify the optimal sports position based on the user's physical characteristics and emotional performance.
[0527] Step 6:
[0528] The server generates position suggestions based on the model's execution results and sends them to the user's terminal. The terminal receives this data and presents it to the user in an easy-to-read format. The user can review these suggestions and use them to adjust their own training strategy.
[0529] (Example 2)
[0530] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0531] In sports, position selection for athletes is often based primarily on physical characteristics, with a lack of comprehensive evaluation and suggestions that consider emotional state and psychological condition. This makes it difficult to find the optimal position that allows athletes to perform at their best.
[0532] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0533] In this invention, the server includes signal acquisition means for inputting video signals acquired from a camera that records the user's movements, feature extraction means for extracting characteristics of the user's movements based on the video signals, and emotion recognition means for identifying emotional states from facial expressions and sounds in the video. This makes it possible to suggest the optimal sports position that takes into account both the user's movement characteristics and emotional state.
[0534] "Signal acquisition means" refers to means of inputting video signals acquired from a camera used to record the user's actions.
[0535] A "feature extraction means" is a means for extracting the characteristics of a user's actions based on a video signal.
[0536] "Emotion recognition means" refers to methods for identifying an emotional state from the user's facial expressions and voice in a video.
[0537] A "means for receiving responses" refers to a method for presenting questions to users and receiving their answers.
[0538] "Evaluation methods" refer to means of evaluating a user's suitability by combining characteristic data, sentiment data, and response data.
[0539] The "proposal method" refers to a means of suggesting the optimal sports position to the user using a generative AI model based on the evaluated aptitude.
[0540] This invention is a comprehensive sports aptitude evaluation system that utilizes video recognition technology and emotion recognition technology, with the aim of supporting users in performing sports in the optimal position. This system is configured as follows.
[0541] First, users record their sports performance in video format using recording devices such as smartphones or video cameras. The recorded videos are saved on the user's device and then uploaded to a server via the network. The server acquires this video signal and extracts action features using video recognition technology. This process uses OpenCV, an open-source video analysis library. The server also analyzes the user's facial expressions and voice in the video using Google Cloud's sentiment analysis API and other tools to recognize their emotional state.
[0542] Next, the user answers questions about sports on their device, and this answer data is sent to the server. This question-and-answer data provides detailed information about the user's self-assessment and playing style, and, combined with emotion and behavior data, forms the basis for further analysis.
[0543] The server integrates this data and uses a generative AI model to evaluate the most suitable sports position for the user. This evaluation process takes the user's emotional data into particular, making it possible to suggest a balanced position that reflects their psychological condition. For example, the server can detect increases in stress from the user's gameplay video and suggest a position that allows the user to relax accordingly.
[0544] The evaluation results are sent from the server to the user's terminal and presented visually. An example of a prompt message here is the command, "Suggest the optimal position based on the motion and emotion data extracted from the user's video." In this way, the system considers both the user's physical characteristics and emotional elements, providing strong support for improving and optimizing playing style in sports.
[0545] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0546] Step 1:
[0547] Users record their sports play in video format using their smartphones or video cameras. The input is the user's actions during play, and the output is the recorded video data. This process involves using the device's camera function to start recording and capturing specific plays.
[0548] Step 2:
[0549] The user uploads recorded video data from their device to the server via the internet. The input is the video data stored on the user's device, and the output is the video being sent to the server. This process involves pressing the "Upload" button within the application and establishing a connection with the server.
[0550] Step 3:
[0551] The server receives uploaded video data and extracts motion features using video recognition technology. The input is video data, and the output is motion feature data. Specifically, this involves analyzing body movements using the OpenCV library and calculating joint positions and movement patterns.
[0552] Step 4:
[0553] The server uses emotion recognition tools to identify the emotional state of the user from their facial expressions and voice in a video. The input is video data, and the output is emotional state data. Specifically, this involves using an emotion analysis API to analyze changes in facial expressions and voice tone to identify the type of emotion.
[0554] Step 5:
[0555] The user answers a sports-related question form on their device and sends the answers to the server. The input is the user's answers, and the output is the response data transferred to the server. The specific actions include entering information into the form and pressing the "Submit" button.
[0556] Step 6:
[0557] The server integrates and analyzes behavioral characteristic data, emotional state data, and user question response data. The input includes this data, and the output is the integrated analysis results. This includes using a generative AI model to analyze this information and evaluate the optimal position for the user.
[0558] Step 7:
[0559] The server sends the evaluation results to the user's terminal, which then displays the results visually. The input includes the analysis results, and the output is the evaluation results displayed on the terminal for the user to view. This process includes a function to display detailed evaluation information by pressing the "View Results" button.
[0560] (Application Example 2)
[0561] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0562] In conventional factory work, it has been difficult to suggest the optimal work position while considering not only the worker's physical movements but also their emotional changes, resulting in decreased work efficiency and increased psychological burden. There is a need for a system that comprehensively evaluates the physical and mental state of workers in the work environment and recommends the most effective and least stressful work position and tasks.
[0563] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0564] In this invention, the server includes signal acquisition means for inputting video signals acquired from a camera for recording user movements, feature extraction means for extracting characteristics of the user's movements based on the video signals, and emotion recognition means for recognizing the user's emotions and generating emotion data. This makes it possible to analyze the user's movements and emotions in real time and propose the optimal work position and tasks.
[0565] A "signal acquisition means" is a device or module that has the function of inputting video signals acquired from a camera into the system.
[0566] A "feature extraction means" is a device or system that analyzes acquired video signals and has the function of extracting the type of user's movements and physical characteristics.
[0567] "Emotion recognition means" refers to devices or systems that have the function of recognizing the emotional state of a user using voice analysis or facial expression analysis and generating emotional data.
[0568] A "response receiving means" refers to a device or module that has the function of presenting questions to users and receiving their responses.
[0569] "Evaluation means" refers to a device or module that has the function of evaluating the user's suitability by combining feature data obtained by feature extraction means, emotion data generated by emotion recognition means, and response data received by response reception means.
[0570] A "suggestion means" is a device or system that has the function of suggesting the optimal work position or task to the user based on the aptitude evaluated by the evaluation means.
[0571] The system for implementing this invention has the function of analyzing the user's actions and emotional state in real time and suggesting the optimal work position and tasks. The server receives video signals transmitted from the user's terminal through signal acquisition means and extracts characteristics of the actions by analyzing them. Furthermore, using emotion recognition means, it recognizes emotions from the user's face and voice and generates emotion data.
[0572] The terminal presents the user with questions about their work and receives their answers through a response reception mechanism. By combining these answers with data obtained by feature extraction and emotion recognition mechanisms, an evaluation mechanism within the server performs a comprehensive suitability assessment, and the suggestion mechanism uses the evaluation results to propose the optimal work position and tasks. This system enables the provision of a more balanced task assignment and work environment that reflects the user's behavior patterns and emotional state.
[0573] For example, consider a situation in a manufacturing plant where a worker is standing for long periods. This system can detect the user's fatigue and stress in real time and, based on that, suggest the optimal position and relatively lighter tasks for them. This makes it possible to reduce the worker's workload while maintaining productivity.
[0574] An example of a prompt message is, "Based on the worker's current emotional state, which task or position is appropriate?" Based on this prompt message, the AI model evaluates and generates suggestions.
[0575] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0576] Step 1:
[0577] The server receives video signals transmitted from the user's terminal using signal acquisition means. The video signal is provided as input, and preprocessing is performed based on it for motion analysis. This preprocessing removes noise from the video and converts it into an analyzable format.
[0578] Step 2:
[0579] Using the processed video data, the server's feature extraction mechanism extracts motion characteristics. Pre-processed video data is used as input, and the user's motion patterns are generated as output. Specifically, the type and frequency of movement, the trajectory of body movements, etc., are analyzed by an algorithm.
[0580] Step 3:
[0581] Simultaneously, the server's emotion recognition system analyzes the user's face and voice in the video to generate emotion data. Voice and facial feature data are used as input. Based on facial expressions and tone of voice, it recognizes emotions such as satisfaction, dissatisfaction, and stress.
[0582] Step 4:
[0583] The terminal presents the user with questions about specific tasks and receives the user's answers through a response reception mechanism. The input includes the answers selected by the user on the terminal and is transmitted directly to the server. The questions often concern the user's experience and level of comfort.
[0584] Step 5:
[0585] The server integrates data obtained by feature extraction and sentiment recognition methods, along with user response data, using evaluation methods to assess the user's suitability. All collected data is included as input, and the output generates recommendations for the most suitable work position and tasks for the user.
[0586] Step 6:
[0587] Based on the evaluation results, the server's suggestion mechanism uses a generated AI model to create prompt messages that suggest the optimal work position and tasks for the user. The output includes specific work location suggestions and instructions for the user, and is presented to the user via the terminal.
[0588] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0589] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0590] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0591] [Fourth Embodiment]
[0592] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0593] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0594] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0595] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0596] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0597] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0598] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0599] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0600] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0601] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0602] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0603] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0604] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0605] This system aims to support users in performing sports from the optimal position, and is a sports aptitude evaluation system that utilizes video recognition technology and machine learning. The embodiments for carrying out the present invention will be described in detail below.
[0606] The user first records their sports play in video format using a recording device. This recorded video is then uploaded from the terminal to a server via the network. The server processes the received video and performs motion analysis. Specifically, the server divides the video into frames and uses image recognition technology to extract the motion of each frame as numerical data. This feature data is useful for detailed analysis of the user's movements, physique, form, etc.
[0607] Simultaneously, users are required to answer questions via their device regarding their sports experience, playing style, and preferred movements. The device then sends these answers to a server. This question-and-answer data integrates the user's subjective information with objective data.
[0608] The server combines feature data and question-answer data, and uses a machine learning model to evaluate the most suitable position for the user based on this information. It compares this information with data on athletes obtained from past databases to assess suitability. This makes it possible to recommend the position best suited to the user's movement style and characteristics.
[0609] The recommended results are sent from the server to the user's device, which then displays them to the user. The user can review the suggested results and use them to practice and train for the newly evaluated position. For example, if the server recommends outfielder as the optimal position based on the user's pitching motion and running ability, the user can use that feedback to refine their suitability as an outfielder.
[0610] In this way, this system maximizes the user's potential and supports the improvement and optimization of their playing style in sports.
[0611] The following describes the processing flow.
[0612] Step 1:
[0613] The user records their sports play with a recording device and saves the video file to their device. Using the device's interface, they upload the recorded video to the server. The server receives the uploaded video file and saves it to its storage.
[0614] Step 2:
[0615] The server divides the saved video into frames and applies an image recognition algorithm to each frame to extract motion characteristics. During this process, the server collects information such as the user's posture, movement speed, and angle as numerical data. The extracted data is temporarily stored.
[0616] Step 3:
[0617] The user is asked to answer questions about their sports experience. The device displays a question form, and the user enters information about their experience, playing style, and preferred movements. The device then sends these answers to the server.
[0618] Step 4:
[0619] The server receives image recognition feature data extracted from the frames and question data answered by the user, and integrates this data. The server then performs preprocessing to prepare the data for use in machine learning models.
[0620] Step 5:
[0621] The server inputs the prepared data into a machine learning model and begins the process of evaluating the user's suitability. The model analyzes influential features and performs calculations to recommend the optimal position by comparing them with past data.
[0622] Step 6:
[0623] The server sends the user the optimal sports position derived from machine learning. The user's device receives this information and displays the results in a visually appealing format. Based on this information, the user adjusts their training and plans how to adapt to the suggested position.
[0624] (Example 1)
[0625] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0626] In today's sports environment, quickly and accurately determining the optimal playing position based on a player's characteristics is a challenging task. Traditional methods often rely on subjective evaluations and lack objective, data-driven analysis, which can prevent players from fully realizing their potential.
[0627] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0628] In this invention, the server includes data acquisition means, feature analysis means, and information gathering means. This makes it possible to integrate and evaluate the user's movement characteristics and subjective information, and objectively recommend the optimal movement position.
[0629] "Data acquisition means" refers to a function that inputs video data acquired from a camera that records the user's actions.
[0630] A "feature analysis method" is a function that divides video data into frames and extracts the actions of each frame as numerical data.
[0631] An "information gathering tool" is a function that presents questions to the user and receives the answers in digital format.
[0632] The "analysis means" is a function that integrates numerical data extracted by the feature analysis means with response data received by the information collection means to evaluate the user's characteristics.
[0633] "Recommendation method" refers to a function that recommends the optimal exercise position to the user based on the evaluation results from the analysis method.
[0634] This invention is a system that analyzes the user's movement characteristics based on data and proposes the optimal sports position. Specific embodiments of this system are shown below.
[0635] System configuration:
[0636] The system primarily consists of a server, terminals, and user operations. Users record their sports play in video format using recording devices (e.g., smartphones or dedicated cameras). Users upload these recorded videos to the server via their terminals. The server divides the received video into frames and uses image recognition technology as a feature analysis tool to extract and analyze the actions in each frame as numerical data. In particular, the use of deep learning libraries and computer vision models is conceivable.
[0637] Data collection:
[0638] Users answer questions about their sports experience and preferred playing style through a terminal. This information is entered into the terminal in digital format and then transmitted to a server. The server aggregates a large amount of question-and-answer data through this information collection method.
[0639] Data integration and analysis:
[0640] The server integrates numerical data extracted by feature analysis tools with question response data collected by information gathering tools. Based on the resulting dataset, machine learning techniques are used as an analytical tool to objectively evaluate the user's characteristics. The generative AI model compares the user with a database of past athletes to recommend the most suitable sports position for the user.
[0641] Presentation of results:
[0642] Based on the evaluation results, the server utilizes recommended methods to present the optimal sports position to the user on the device. The device can then display feedback to the user and provide specific training guidelines.
[0643] Specific example:
[0644] For example, if the server analyzes a user's video data and recommends that "outfielder" is the best position based on their pitching motion and excellent running ability, the user can then train based on this evaluation.
[0645] Example prompts for generative AI models:
[0646] "Evaluate the user's pitching and running abilities based on their video data, and recommend a suitable sports position."
[0647] "Consider your experience and playing style, and select the position that best suits you."
[0648] This invention provides a precise approach to maximizing sports performance.
[0649] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0650] Step 1:
[0651] The user records their sports play in video format using a recording device. The input is the user's actions, and video data is generated as output. Specifically, the user uses a smartphone to film their soccer dribbling movements.
[0652] Step 2:
[0653] The device uploads video data to the server. The input is video data, and the output is data transfer to the server via the network. For example, pressing the "upload" button in the dedicated application sends the video to the server.
[0654] Step 3:
[0655] The server divides the received video into frames. The input is the video data transferred to the server, and the output is a dataset of image frames. Specifically, the server divides the video into 30 frames per second and generates a still image for each frame.
[0656] Step 4:
[0657] The server analyzes each frame of the video and uses image recognition technology as a feature analysis tool to extract motion characteristics. The input is image frame data, and the output is numerical data representing the motion of each frame. Specifically, the server detects things like the position of the feet and the angle of the body, and quantifies these.
[0658] Step 5:
[0659] Users answer questions about their sports experience and preferred playing style through their device. The input is user-provided data, and structured response information is generated as output. Specifically, users answer questions in a survey format such as, "What are your preferred movements?"
[0660] Step 6:
[0661] The terminal sends the user's response data to the server. The input is the user's response information, and the output is the transfer to the server. Specifically, the user presses the "Send" button on the terminal, and the data is sent to the server.
[0662] Step 7:
[0663] The server integrates numerical data extracted by the feature analysis method with response data. The input consists of quantified behavioral data and response data, and the output is an integrated dataset. This data integration enables comprehensive analysis.
[0664] Step 8:
[0665] The server uses integrated data to perform analysis with a generative AI model and evaluate the user's characteristics. The input is the integrated dataset, and the output is a proposal as an evaluation result. Specifically, the server analyzes behavioral patterns using a machine learning algorithm and recommends the optimal position.
[0666] Step 9:
[0667] The server sends the results to the terminal, which then displays the optimal position to the user. The input is the evaluation result sent from the server, and the output is the feedback displayed to the user. The user can then use this feedback to train.
[0668] (Application Example 1)
[0669] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0670] There is a need to improve the efficiency of robots that perform multiple tasks in factories. However, currently, it is difficult to understand the optimal work position and role for each robot, making it challenging to place them in the right place. This leads to decreased productivity and wasted resources.
[0671] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0672] In this invention, the server includes data acquisition means for inputting video signals obtained from a camera for recording user movements, data analysis means for extracting characteristics of the user's movements based on the video signals, and information receiving means for presenting questions to the user and receiving their answers. This makes it possible to recommend the optimal role and position for each robot to perform its work most efficiently.
[0673] The "data acquisition means" is responsible for acquiring video signals from devices that record the movements of users or robots, and transmitting them to a server for analysis.
[0674] A "data analysis means" is a device that uses a visual data recognition algorithm to extract motion characteristics from acquired video signals and analyzes them as numerical data.
[0675] An "information reception method" is a system that presents questions to users or administrators, receives their answers, and transmits them to a server as data for use in suitability assessments.
[0676] "Suitability evaluation means" refers to a system that integrates feature data obtained from data analysis means and information data acquired from information reception means to evaluate the most suitable work position and role for each robot.
[0677] The "role proposal method" is a process that proposes the most suitable work role for a robot based on the evaluation results derived from the aptitude evaluation method, and supports improvements in placement and work content.
[0678] This system is designed to optimize the work efficiency of robots in factories. First, cameras are attached to the robots in the factory, and these cameras continuously record the robots' movements. This serves as a means of data acquisition.
[0679] The server uses video recognition technologies such as OpenCV to divide the acquired video signal into frames and processes the motion of each frame using data analysis tools. This process extracts the robot's motion characteristics and external features as numerical data. Furthermore, the server uses machine learning frameworks such as TensorFlow and PyTorch to analyze the data.
[0680] As a means of receiving information, administrators are provided with tablets or PCs, on which questions from the server are presented. Administrators answer these questions, and the content of their answers is sent to the server. The server uses these answers to evaluate the robot's suitability using an aptitude evaluation system.
[0681] As a result, the optimal robot work position and role are suggested by the role suggestion system. This information can be viewed on the factory's management system.
[0682] For example, if a suitability assessment recommends that a particular robot be specialized for transport tasks in order to improve the efficiency of assembly and transport operations, the factory manager can change the robot's placement based on that suggestion.
[0683] An example of a prompt in a generative AI model is, "Based on the operation data of the factory robot, please suggest the most efficient work role."
[0684] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0685] Step 1:
[0686] A factory robot equipped with a camera records its own movements as video during work. The input is real-time video data. This video data is later uploaded to a server as a data acquisition method.
[0687] Step 2:
[0688] The server divides the uploaded video data into frames. The input is the video data, and the output is a collection of still images divided into each frame. The server uses these divided frames to prepare for extracting the operating characteristics using OpenCV.
[0689] Step 3:
[0690] The server uses OpenCV as a data analysis tool to extract numerical data representing the robot's motion characteristics from each frame. The input is the divided frames, and the output is numerical data showing the motion characteristics. The server then uses this numerical data for subsequent analysis.
[0691] Step 4:
[0692] The server analyzes numerical data using machine learning frameworks such as TensorFlow or PyTorch. The input is numerical data representing operational characteristics, and the output is analysis results related to the suitability evaluation of the robots. Based on these analysis results, the server evaluates the efficiency of each robot.
[0693] Step 5:
[0694] The terminal (administrator's tablet or PC) functions as hardware that displays questions sent from the server and prompts the administrator to answer. The input is question data from the server, and the output is answer data from the administrator. The specific actions performed on the terminal are the presentation of questions and the acceptance of answers.
[0695] Step 6:
[0696] The server uses the response data to the questions to perform an aptitude assessment. The input consists of the analysis results of operational characteristics and the response data, and the output is the assessed aptitude data. Based on this data, the server calculates the optimal role.
[0697] Step 7:
[0698] The server utilizes a generated AI model based on the suitability assessment results to propose the most efficient robot role. The input is suitability data, and the output is information about the proposed role. The server transmits this information to the factory management system.
[0699] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0700] The present invention is a comprehensive sports aptitude evaluation system that utilizes video recognition technology and emotion recognition technology, with the aim of supporting users in performing sports in the optimal position. The embodiments for carrying out the present invention will be described in detail below.
[0701] The user first records their sports play in video format using a recording device. This video is uploaded from the user's device to a server via the network. The server analyzes the received video as described above to extract characteristics of the user's movements. Furthermore, using an emotion engine, it simultaneously recognizes changes in the user's emotions through facial expressions and voice from the video. This allows for the evaluation to incorporate not only movements but also the emotional aspects of play.
[0702] In parallel, users answer questions about sports. This interaction takes place on the device, and detailed answers from the user are sent to the server via a question form. These answers provide information about the user's self-assessment and playing style, and, combined with emotional and behavioral data, form the basis for further analysis.
[0703] The server combines behavioral feature data, emotional state data, and question response data, and uses a machine learning model to evaluate the most suitable position for the user. This evaluation process particularly considers emotional data, analyzing which position elicits the most positive emotions from the user. This enables balanced position recommendations that also take into account the user's psychological condition.
[0704] The evaluation results are sent from the server to the user's terminal, which then visually displays them to the user. For example, the server can detect increases in the user's stress levels from their gameplay video and suggest a more relaxed playing position accordingly, allowing the user to concentrate better and perform more effectively.
[0705] In this way, the system takes into account both the user's physical characteristics and emotional factors, providing strong support for improving and optimizing their playing style in sports.
[0706] The following describes the processing flow.
[0707] Step 1:
[0708] The user records their sports play with a recording device and saves the video data to their device. Next, the user uses the device's interface to upload the recorded video to the system's server. The server saves the received video to its storage.
[0709] Step 2:
[0710] The server processes the saved video frame by frame and extracts motion features using an image recognition algorithm. In this process, the server extracts motion data for each frame (e.g., speed of movement, body angle) as numerical data and temporarily stores it in a database.
[0711] Step 3:
[0712] The server then activates an emotion engine that analyzes the user's facial expressions and voice tone from the video data to recognize the user's emotional state. The recognized emotions are extracted as indicators of stress levels and concentration during gameplay and stored along with the action data.
[0713] Step 4:
[0714] Users answer questions about their sports experience and playing style. The device displays an input form for the questions, where users provide answers through comments and multiple-choice options. The device sends these answers to the server, which adds them to a database.
[0715] Step 5:
[0716] The server integrates behavioral feature data, emotional data, and user response data, and inputs them into a machine learning model to assess the user's suitability. Here, the server uses the model to analyze each dataset and identify the optimal sports position based on the user's physical characteristics and emotional performance.
[0717] Step 6:
[0718] The server generates position suggestions based on the model's execution results and sends them to the user's terminal. The terminal receives this data and presents it to the user in an easy-to-read format. The user can review these suggestions and use them to adjust their own training strategy.
[0719] (Example 2)
[0720] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0721] In sports, position selection for athletes is often based primarily on physical characteristics, with a lack of comprehensive evaluation and suggestions that consider emotional state and psychological condition. This makes it difficult to find the optimal position that allows athletes to perform at their best.
[0722] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0723] In this invention, the server includes signal acquisition means for inputting video signals acquired from a camera that records the user's movements, feature extraction means for extracting characteristics of the user's movements based on the video signals, and emotion recognition means for identifying emotional states from facial expressions and sounds in the video. This makes it possible to suggest the optimal sports position that takes into account both the user's movement characteristics and emotional state.
[0724] "Signal acquisition means" refers to means of inputting video signals acquired from a camera used to record the user's actions.
[0725] A "feature extraction means" is a means for extracting the characteristics of a user's actions based on a video signal.
[0726] "Emotion recognition means" refers to methods for identifying an emotional state from the user's facial expressions and voice in a video.
[0727] A "means for receiving responses" refers to a method for presenting questions to users and receiving their answers.
[0728] "Evaluation methods" refer to means of evaluating a user's suitability by combining characteristic data, sentiment data, and response data.
[0729] The "proposal method" refers to a means of suggesting the optimal sports position to the user using a generative AI model based on the evaluated aptitude.
[0730] This invention is a comprehensive sports aptitude evaluation system that utilizes video recognition technology and emotion recognition technology, with the aim of supporting users in performing sports in the optimal position. This system is configured as follows.
[0731] First, users record their sports performance in video format using recording devices such as smartphones or video cameras. The recorded videos are saved on the user's device and then uploaded to a server via the network. The server acquires this video signal and extracts action features using video recognition technology. This process uses OpenCV, an open-source video analysis library. The server also analyzes the user's facial expressions and voice in the video using Google Cloud's sentiment analysis API and other tools to recognize their emotional state.
[0732] Next, the user answers questions about sports on their device, and this answer data is sent to the server. This question-and-answer data provides detailed information about the user's self-assessment and playing style, and, combined with emotion and behavior data, forms the basis for further analysis.
[0733] The server integrates this data and uses a generative AI model to evaluate the most suitable sports position for the user. This evaluation process takes the user's emotional data into particular, making it possible to suggest a balanced position that reflects their psychological condition. For example, the server can detect increases in stress from the user's gameplay video and suggest a position that allows the user to relax accordingly.
[0734] The evaluation results are sent from the server to the user's terminal and presented visually. An example of a prompt message here is the command, "Suggest the optimal position based on the motion and emotion data extracted from the user's video." In this way, the system considers both the user's physical characteristics and emotional elements, providing strong support for improving and optimizing playing style in sports.
[0735] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0736] Step 1:
[0737] Users record their sports play in video format using their smartphones or video cameras. The input is the user's actions during play, and the output is the recorded video data. This process involves using the device's camera function to start recording and capturing specific plays.
[0738] Step 2:
[0739] The user uploads recorded video data from their device to the server via the internet. The input is the video data stored on the user's device, and the output is the video being sent to the server. This process involves pressing the "Upload" button within the application and establishing a connection with the server.
[0740] Step 3:
[0741] The server receives uploaded video data and extracts motion features using video recognition technology. The input is video data, and the output is motion feature data. Specifically, this involves analyzing body movements using the OpenCV library and calculating joint positions and movement patterns.
[0742] Step 4:
[0743] The server uses emotion recognition tools to identify the emotional state of the user from their facial expressions and voice in a video. The input is video data, and the output is emotional state data. Specifically, this involves using an emotion analysis API to analyze changes in facial expressions and voice tone to identify the type of emotion.
[0744] Step 5:
[0745] The user answers a sports-related question form on their device and sends the answers to the server. The input is the user's answers, and the output is the response data transferred to the server. The specific actions include entering information into the form and pressing the "Submit" button.
[0746] Step 6:
[0747] The server integrates and analyzes behavioral characteristic data, emotional state data, and user question response data. The input includes this data, and the output is the integrated analysis results. This includes using a generative AI model to analyze this information and evaluate the optimal position for the user.
[0748] Step 7:
[0749] The server sends the evaluation results to the user's terminal, which then displays the results visually. The input includes the analysis results, and the output is the evaluation results displayed on the terminal for the user to view. This process includes a function to display detailed evaluation information by pressing the "View Results" button.
[0750] (Application Example 2)
[0751] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0752] In conventional factory work, it has been difficult to suggest the optimal work position while considering not only the worker's physical movements but also their emotional changes, resulting in decreased work efficiency and increased psychological burden. There is a need for a system that comprehensively evaluates the physical and mental state of workers in the work environment and recommends the most effective and least stressful work position and tasks.
[0753] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0754] In this invention, the server includes signal acquisition means for inputting video signals acquired from a camera for recording user movements, feature extraction means for extracting characteristics of the user's movements based on the video signals, and emotion recognition means for recognizing the user's emotions and generating emotion data. This makes it possible to analyze the user's movements and emotions in real time and propose the optimal work position and tasks.
[0755] A "signal acquisition means" is a device or module that has the function of inputting video signals acquired from a camera into the system.
[0756] A "feature extraction means" is a device or system that analyzes acquired video signals and has the function of extracting the type of user's movements and physical characteristics.
[0757] "Emotion recognition means" refers to devices or systems that have the function of recognizing the emotional state of a user using voice analysis or facial expression analysis and generating emotional data.
[0758] A "response receiving means" refers to a device or module that has the function of presenting questions to users and receiving their responses.
[0759] "Evaluation means" refers to a device or module that has the function of evaluating the user's suitability by combining feature data obtained by feature extraction means, emotion data generated by emotion recognition means, and response data received by response reception means.
[0760] A "suggestion means" is a device or system that has the function of suggesting the optimal work position or task to the user based on the aptitude evaluated by the evaluation means.
[0761] The system for implementing this invention has the function of analyzing the user's actions and emotional state in real time and suggesting the optimal work position and tasks. The server receives video signals transmitted from the user's terminal through signal acquisition means and extracts characteristics of the actions by analyzing them. Furthermore, using emotion recognition means, it recognizes emotions from the user's face and voice and generates emotion data.
[0762] The terminal presents the user with questions about their work and receives their answers through a response reception mechanism. By combining these answers with data obtained by feature extraction and emotion recognition mechanisms, an evaluation mechanism within the server performs a comprehensive suitability assessment, and the suggestion mechanism uses the evaluation results to propose the optimal work position and tasks. This system enables the provision of a more balanced task assignment and work environment that reflects the user's behavior patterns and emotional state.
[0763] For example, consider a situation in a manufacturing plant where a worker is standing for long periods. This system can detect the user's fatigue and stress in real time and, based on that, suggest the optimal position and relatively lighter tasks for them. This makes it possible to reduce the worker's workload while maintaining productivity.
[0764] An example of a prompt message is, "Based on the worker's current emotional state, which task or position is appropriate?" Based on this prompt message, the AI model evaluates and generates suggestions.
[0765] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0766] Step 1:
[0767] The server receives video signals transmitted from the user's terminal using signal acquisition means. The video signal is provided as input, and preprocessing is performed based on it for motion analysis. This preprocessing removes noise from the video and converts it into an analyzable format.
[0768] Step 2:
[0769] Using the processed video data, the server's feature extraction mechanism extracts motion characteristics. Pre-processed video data is used as input, and the user's motion patterns are generated as output. Specifically, the type and frequency of movement, the trajectory of body movements, etc., are analyzed by an algorithm.
[0770] Step 3:
[0771] Simultaneously, the server's emotion recognition system analyzes the user's face and voice in the video to generate emotion data. Voice and facial feature data are used as input. Based on facial expressions and tone of voice, it recognizes emotions such as satisfaction, dissatisfaction, and stress.
[0772] Step 4:
[0773] The terminal presents the user with questions about specific tasks and receives the user's answers through a response reception mechanism. The input includes the answers selected by the user on the terminal and is transmitted directly to the server. The questions often concern the user's experience and level of comfort.
[0774] Step 5:
[0775] The server integrates data obtained by feature extraction and sentiment recognition methods, along with user response data, using evaluation methods to assess the user's suitability. All collected data is included as input, and the output generates recommendations for the most suitable work position and tasks for the user.
[0776] Step 6:
[0777] Based on the evaluation results, the server's suggestion mechanism uses a generated AI model to create prompt messages that suggest the optimal work position and tasks for the user. The output includes specific work location suggestions and instructions for the user, and is presented to the user via the terminal.
[0778] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0779] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0780] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0781] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0782] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0783] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0784] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0785] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0786] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0787] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0788] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0789] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0790] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0791] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0792] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0793] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0794] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0795] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0796] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0797] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0798] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0799] The following is further disclosed regarding the embodiments described above.
[0800] (Claim 1)
[0801] A signal acquisition means that receives video signals acquired from a camera for recording the user's movements,
[0802] A feature extraction means for extracting the characteristics of the user's actions based on the aforementioned video signal,
[0803] A means of receiving answers by presenting questions to users and receiving their responses,
[0804] An evaluation means for evaluating the user's suitability by combining the feature data extracted by the feature extraction means and the response data received by the response receiving means,
[0805] Based on the aptitude evaluated by the aforementioned evaluation means, a suggestion means proposes the optimal sports position for the user,
[0806] A system that includes this.
[0807] (Claim 2)
[0808] The system according to claim 1, characterized in that the feature extraction means has a function to identify the type of movement and body size using an image recognition algorithm.
[0809] (Claim 3)
[0810] The system according to claim 1, characterized in that the proposed means includes a function that recommends a position while comparing it with a past database using a machine learning model.
[0811] "Example 1"
[0812] (Claim 1)
[0813] A data acquisition means that inputs video data acquired from a camera for recording the user's movements,
[0814] The aforementioned video data is divided into frames, and a feature analysis means extracts the actions of each frame as numerical data,
[0815] A means of gathering information that presents questions to users and receives their answers in digital format,
[0816] An analysis means that integrates the numerical data extracted by the feature analysis means and the response data received by the information collection means to evaluate the characteristics of the user,
[0817] Based on the evaluation results from the aforementioned analysis means, a recommendation means recommends the optimal exercise position for the user,
[0818] A system that includes this.
[0819] (Claim 2)
[0820] The system according to claim 1, characterized in that the feature analysis means includes a function to identify indicators of movement and physical characteristics using image recognition technology.
[0821] (Claim 3)
[0822] The system according to claim 1, characterized in that the recommendation means includes a function that uses machine learning technology to recommend the optimal position while comparing it with past databases.
[0823] "Application Example 1"
[0824] (Claim 1)
[0825] A data acquisition means that inputs video signals acquired from a camera for recording user movements,
[0826] A data analysis means for extracting the characteristics of the user's movements based on the aforementioned video signal,
[0827] An information receiving method that presents questions to users and receives their answers,
[0828] An aptitude evaluation means that evaluates the user's suitability by combining the feature data extracted by the data analysis means and the information data received by the information receiving means,
[0829] A role suggestion means that proposes the most suitable work role to the user based on the aptitude evaluated by the aforementioned aptitude evaluation means,
[0830] A system that includes this.
[0831] (Claim 2)
[0832] The system according to claim 1, characterized in that the data analysis means has a function to identify the type of action and appearance characteristics using a visual data recognition algorithm.
[0833] (Claim 3)
[0834] The system according to claim 1, characterized in that the role suggestion means has a function that recommends roles using a machine learning model while comparing them with a past information base.
[0835] "Example 2 of combining an emotion engine"
[0836] (Claim 1)
[0837] A signal acquisition means that receives video signals acquired from a camera for recording the user's movements,
[0838] A feature extraction means for extracting the characteristics of the user's actions based on the aforementioned video signal,
[0839] An emotion recognition method that identifies the emotional state from the user's facial expressions and voice in the video,
[0840] A means of receiving answers by presenting questions to users and receiving their responses,
[0841] An evaluation means for evaluating the user's suitability by combining the data processed by the feature extraction means and the emotion recognition means and the response data received by the response receiving means,
[0842] A proposal means that uses a generative AI model to suggest the optimal sports position for the user based on the aptitude evaluated by the evaluation means,
[0843] A system that includes this.
[0844] (Claim 2)
[0845] The system according to claim 1, characterized in that the feature extraction means has a function to identify the type of movement and physique using an image recognition algorithm, as well as a function to analyze changes in facial expression using an emotion engine.
[0846] (Claim 3)
[0847] The system according to claim 1, characterized in that the proposed means includes a function that uses a machine learning model to recommend a sports position while taking into particular consideration user sentiment data and comparing it with a past database.
[0848] "Application example 2 when combining with an emotional engine"
[0849] (Claim 1)
[0850] A signal acquisition means that receives video signals acquired from a camera for recording the user's movements,
[0851] A feature extraction means for extracting the characteristics of the user's actions based on the aforementioned video signal,
[0852] An emotion recognition means that recognizes the user's emotions and generates emotion data,
[0853] A means of receiving answers by presenting questions to users and receiving their responses,
[0854] An evaluation means for evaluating the user's suitability by combining the feature data extracted by the feature extraction means, the emotion data generated by the emotion recognition means, and the response data received by the response receiving means,
[0855] A proposal means that suggests the optimal work position for the user based on the aptitude evaluated by the evaluation means,
[0856] A system that includes this.
[0857] (Claim 2)
[0858] The system according to claim 1, characterized in that the feature extraction means has a function to identify the type of movement and physique using an image recognition algorithm, and the emotion recognition means has a function to evaluate emotions using voice analysis and facial expression analysis.
[0859] (Claim 3)
[0860] The system according to claim 1, characterized in that the proposed means has a function that uses a machine learning model to recommend a position while comparing it with a past database and to suggest an optimal work task based on the user's emotions and behavioral patterns. [Explanation of Symbols]
[0861] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A signal acquisition means that receives video signals acquired from a camera for recording the user's movements, A feature extraction means for extracting the characteristics of the user's actions based on the aforementioned video signal, A means of receiving answers by presenting questions to users and receiving their responses, An evaluation means for evaluating the user's suitability by combining the feature data extracted by the feature extraction means and the response data received by the response receiving means, Based on the aptitude evaluated by the aforementioned evaluation means, a suggestion means proposes the optimal sports position for the user, A system that includes this.
2. The system according to claim 1, characterized in that the feature extraction means has a function to identify the type of movement and body size using an image recognition algorithm.
3. The system according to claim 1, characterized in that the proposed means includes a function that recommends a position using a machine learning model while comparing it with a past database.