system
A system that extracts skeletal characteristics from image data to recommend suitable sports for children, addressing the challenge of parental confusion by offering objective and secure sport recommendations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-25
AI Technical Summary
Parents struggle to determine a suitable sport for their children without appropriate information and expertise, leading to confusion due to the abundance of options.
A system that acquires image data to extract skeletal characteristics, compares them with pre-stored athlete data, and recommends suitable sports based on scientific and objective grounds while ensuring privacy through encryption.
Enables users to make rational, scientifically-based sport selections for children, providing reliable information and protecting privacy.
Smart Images

Figure 2026085780000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Finding a suitable sport based on a child's skeleton poses a problem that it is difficult without appropriate information and expertise. Parents are often confused by the abundance of options and have difficulty determining which sport is suitable for their child. In such a situation, there is a need for a method or system that can easily identify a sport suitable for a child based on scientific and objective grounds based on the skeleton.
Means for Solving the Problems
[0005] This invention provides a means for acquiring image data and extracting a child's skeletal characteristics from said image data. By comparing the extracted skeletal characteristics with pre-stored athlete data, a system is constructed to determine a suitable sport. This system is easy for users to use and suggests the sport that best matches the child's skeletal structure. Furthermore, by encrypting the determination results and enabling secure communication, the system provides users with reliable information while protecting their privacy.
[0006] "Image data" refers to visual information expressed in digital format, including photographs, illustrations, and other digital media.
[0007] "Skeletal characteristics" refer to information about specific shapes and dimensions in the body's structure, and mainly include physical characteristics such as height, limb length, and shoulder width.
[0008] "Athlete data" refers to information about past and present athletes, including detailed data such as skeletal structure, physical abilities, and athletic performance.
[0009] "Encryption" is a type of security measure that transforms information using special methods to prevent third parties from understanding its contents.
[0010] A "user" is an individual or group that uses the system, and in this invention, it particularly refers to parents or guardians who want to find a sport that is suitable for their child. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4]This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0013] First, let's explain the terminology used in the following explanation.
[0014] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0015] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0017] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] [First Embodiment]
[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] To implement this invention, a user acquires image data of a child's entire body using a dedicated application. Upon receiving this image data, the terminal first performs preprocessing, such as noise reduction and size adjustment. Subsequently, it extracts skeletal features from the image using computer vision technology. Specifically, it identifies the positions of the face, shoulders, elbows, knees, and ankles, and calculates their relative distances and angles.
[0033] The extracted skeletal features are sent to the server in a standardized format. Upon receiving these skeletal features, the server compares them to a database of athletes and applies a machine learning algorithm to determine the most suitable sport. The server refers to the profiles of numerous athletes stored in the database and identifies sports with statistically similar skeletal characteristics. Based on this analysis, it generates a list of the most suitable sports and the reasons for their selection, and sends it to the user's terminal.
[0034] For example, suppose an analysis of image data of a specific child reveals that the child is tall and has long arms. The server determines that these characteristics are highly correlated with sports where they are typically advantageous, such as basketball, volleyball, and certain track and field events. It then generates a recommendation list that includes these sports and sends it to the terminal with an explanation such as "height and arm length are related."
[0035] In this way, the system of the present invention provides a means to enable users to make scientifically and objectively rational choices.
[0036] Based on the information received, users can assess their child's aptitude and create a concrete plan for sports activities. This helps unlock the child's potential and provides them with opportunities to enjoy sports in an optimal environment.
[0037] The following describes the processing flow.
[0038] Step 1:
[0039] The user uploads image data of the child's entire body to a dedicated application. The user ensures that the child is sitting upright when the photo is taken.
[0040] Step 2:
[0041] The device preprocesses the image data it receives. Specifically, it adjusts the image resolution and removes noise as needed to prepare the image for analysis.
[0042] Step 3:
[0043] The device uses a computer vision algorithm to extract skeletal features from image data. The algorithm identifies the location of each point in the skeleton, such as the shoulders, elbows, knees, and ankles, and calculates the distances and angles between these locations.
[0044] Step 4:
[0045] The terminal converts the extracted skeletal features into a standardized format, encrypts them for security, and then sends them to the server.
[0046] Step 5:
[0047] The server analyzes the skeletal features it receives and compares them to accumulated athlete data. Here, a machine learning algorithm is used to identify the sport with the highest correlation.
[0048] Step 6:
[0049] Based on the analysis results, the server generates a list of the most suitable sports for the child, along with the reasons for their selection. This includes explaining why certain skeletal characteristics are advantageous for specific sports.
[0050] Step 7:
[0051] The server encrypts the generated sports recommendation list and sends it to the terminal. Privacy protection is taken into consideration during this process.
[0052] Step 8:
[0053] The device decrypts the received data and displays it in an easy-to-understand format for the user. Based on this, the user can then plan their child's sports activities.
[0054] (Example 1)
[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0056] Conventional sports aptitude assessment systems suffered from insufficient preprocessing, data transmission, and analysis methods in the process from image data acquisition to aptitude assessment, resulting in challenges in terms of prediction accuracy and security. Furthermore, the feedback provided to users before suggesting appropriate exercise activities was inadequate, sometimes failing to deliver reliable judgments.
[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0058] In this invention, the server includes means for preprocessing image information, means for transmitting joint features in a standard format, and means for analyzing the data using a clustering algorithm. This enables effective and reliable assessment of suitability for exercise activities by statistically comparing it with data from athletes.
[0059] "Image information" refers to visual data acquired in digital format that includes the visual characteristics of a specific object.
[0060] "Preprocessing" refers to the process of applying noise reduction, size adjustment, and other adjustments to acquired image information to prepare the data for a useful format.
[0061] "Joint features" refer to data extracted from image information that shows the location of joints in the human body and their spatial relationships.
[0062] A "standard format" is a structured data format that is consistently used in the transmission, reception, and processing of data.
[0063] A "clustering algorithm" refers to a computational method used to statistically analyze data and group elements that have similar characteristics.
[0064] "Athlete data" refers to a collection of information regarding the physical characteristics and performance of athletes engaged in diverse athletic activities.
[0065] "Assessing suitability for physical activity" is the act of identifying the most suitable exercise or sport for an individual based on analyzed data.
[0066] To implement this invention, the user first uses a dedicated application to acquire full-body image information of a child using a camera-equipped device. The terminal then performs preprocessing on the received image information, such as noise reduction and size adjustment, using an image processing library. Specifically, the terminal utilizes image processing software such as OpenCV to optimize the image quality.
[0067] Next, the device utilizes computer vision technology to extract joint features from the image. This process uses OpenPose, a pre-trained model, to calculate the coordinates of the face and joint positions. This information is structured in a standard format such as JSON and treated as data.
[0068] The terminal sends standardized joint feature data to the server using the HTTPS protocol. After receiving this data, the server applies a clustering algorithm to compare it with the accumulated data on athletes. The server uses machine learning libraries such as Scikit-learn and TENSORFLOW® to identify sports with similar features to the input data, employing techniques such as K-means clustering.
[0069] Based on this analysis, the server generates a list of suitable sports and the reasons for their selection, and sends the information back to the user's device. The user can then review these results on their device and plan exercise activities that are tailored to their child's characteristics.
[0070] For example, if an analysis of a child's image data reveals features such as being tall and having long arms, the server will determine that these features are advantageous in sports like basketball or volleyball. This determination is then returned to the user with a reason for the selection, such as "because height and arm length are correlated."
[0071] An example of a prompt statement can be expressed as follows:
[0072] "Please provide information on a system that suggests optimal exercise activities based on joint features identified from images of children. Specifically, I'd like to know which features correlate with which sports, and how the server analyzes and returns the results."
[0073] Thus, the system of the present invention provides a means to support sports selection in a scientific and objective manner.
[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0075] Step 1:
[0076] The user uses a dedicated application to capture full-body images of their child using the camera on their smartphone or tablet. The images are saved in JPEG or PNG format. The captured image information is then transmitted to the device through the application's interface.
[0077] Step 2:
[0078] The terminal performs preprocessing on the received image information. It receives an image in JPEG or PNG format as input, uses an image processing library to remove noise (using a Gaussian filter) and adjust the size (resizing to 224x224 pixels), and generates a clean, appropriately sized image as output.
[0079] Step 3:
[0080] The device extracts joint features using pre-processed images. Using pre-processed images as input, it detects the positions of faces and joints using computer vision techniques such as OpenPose. As output, it generates a dataset containing coordinate information for each joint.
[0081] Step 4:
[0082] The terminal converts the extracted joint features into a standard format (such as JSON). It uses joint coordinate information as input and converts this data into a structured format such as JSON. As output, it generates a data package that can be sent to the server.
[0083] Step 5:
[0084] The terminal sends a data package of joint features to the server. Using the HTTPS protocol, the data is securely transferred to the server, and once the transmission is complete, the terminal enters a state of waiting for a response from the server.
[0085] Step 6:
[0086] The server performs analysis based on the received joint feature data. Using the received coordinate data as input, it analyzes the data using algorithms such as K-means clustering, utilizing machine learning libraries (Scikit-learn, TensorFlow, etc.). As output, it generates a list of exercise activities suitable for the subject.
[0087] Step 7:
[0088] The server generates analysis results (a list of suitable exercise activities and the reasons for them), encrypts them, and sends them to the terminal. The output provides data in a format viewable on the user's terminal.
[0089] Step 8:
[0090] Users receive feedback from the server on their device and review a list of suitable sports. Based on this information, they can plan their child's sports activities.
[0091] (Application Example 1)
[0092] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0093] In delivery services, a challenge is that delivery personnel are unable to select the most efficient delivery areas and routes based on their individual physical characteristics. As a result, they may not be able to make the most of their individual abilities, potentially leading to decreased work efficiency. This invention aims to eliminate this inefficiency and provide a system that proposes an appropriate delivery process for delivery personnel.
[0094] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0095] In this invention, the server includes means for acquiring image data, means for extracting physical characteristics from the image data, means for comparing the physical characteristics with pre-stored individual data to determine the optimal behavioral pattern, means for presenting the determination result to the user, and means for recommending an efficient region or route based on the determination result. This makes it possible to suggest the optimal delivery region and route for each delivery person.
[0096] "Image data" refers to a record of the visual information of an object in numerical format.
[0097] "Physical characteristics" refer to data that describes an individual's skeletal structure and physical features.
[0098] "Individual data" refers to a collection of information that accumulates the physical characteristics of multiple individuals.
[0099] A "behavioral pattern" is a pattern of optimal behavior or activity that is predicted based on an individual's characteristics.
[0100] "Users" refer to entities or individuals that utilize this system.
[0101] "Region or route" refers to the specific locations or travel routes where the delivery person performs their duties.
[0102] "Recommendation" means presenting the option that best suits specific conditions.
[0103] In implementing this invention, the user first uses a terminal such as a smartphone to acquire whole-body image data of the target individual. The terminal utilizes advanced image processing techniques to extract physical characteristics from the image data. For this processing, OpenCV, an open-source computer vision library, is used for noise reduction and size adjustment.
[0104] The physical characteristics data extracted by the device is sent to the server. The server compares the received physical characteristics with existing data in the individual database and uses a generative AI model to determine the optimal behavioral pattern. During this process, the server performs data analysis using machine learning libraries such as TensorFlow and PyTorch.
[0105] The results are presented to the user. These results include the most suitable region or route for the individual, and suggest efficient actions and activities.
[0106] For example, if a delivery person has physical characteristics such as being tall and having long arms, the server will determine that they are suitable for long-distance deliveries and suggest a delivery route that covers a wide area. This suggestion is then notified to the user's smartphone.
[0107] An example of a prompt message to be input into the generating AI model is: "Generate an efficient delivery route based on the physical characteristics of the delivery person. For example, for a tall delivery person with long arms, suggest a route that covers a wide area." In this way, the present invention provides a system that enables efficient work suited to the characteristics of the delivery person.
[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0109] Step 1:
[0110] The user uses their smartphone camera to take a full-body image of the delivery person. The input is the captured full-body image data, and the output is an image file stored on the device. The device starts processing this image data as its initial input.
[0111] Step 2:
[0112] The terminal uses the OpenCV library to preprocess the acquired image data. The input is whole-body image data, and filtering and resizing are performed to remove noise and adjust the size of the image. The output is preprocessed, clear image data. Specifically, the image sharpness is improved and the resolution is adjusted to facilitate analysis.
[0113] Step 3:
[0114] The device uses TensorFlow to extract body characteristics based on pre-processed image data. The input is pre-processed image data, and the output is body characteristic data such as the position and length of the skeleton. This data extraction uses a pre-trained model to identify the positions of the face, shoulders, elbows, knees, and ankles, and to calculate their relative distances and angles.
[0115] Step 4:
[0116] The terminal sends extracted physical characteristic data to the server. The server receives this data and compares it with existing data in the individual database. The input is physical characteristic data, and the output is identification information for similar database entries. Using a database search algorithm, the server applies a statistical model to efficiently determine similarity.
[0117] Step 5:
[0118] The server runs a machine learning model and compares it with an individual database to determine the optimal behavioral pattern. The input is the physical characteristics data received by the server and the information in the database, and the output is a suggestion of the optimal region or route specific to the individual. This model uses a deep learning algorithm implemented in TensorFlow or PyTorch.
[0119] Step 6:
[0120] The server sends the determined suggestion to the user. The input is the region or route information determined by the server, and the output is a notification message displayed on the user's device. The notification includes a brief explanation of the delivery route details and efficiency. Specifically, the notification message appears as a pop-up on the user's smartphone.
[0121] Step 7:
[0122] Users review notifications and perform delivery tasks according to the suggested area or route. The input is the information presented as a notification, and the output is the streamlined actual delivery task. By adhering to these instructions, users can expect improved operational efficiency and minimized effort.
[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0124] This invention begins with a user using a dedicated application to acquire image data of a child's entire body and uploading this data to a terminal. The image data is preprocessed on the terminal and processed by a computer vision algorithm to extract skeletal features. After detecting the skeletal features, the terminal sends the feature data to a server.
[0125] The server uses the received skeletal features to compare them with pre-stored athlete data and determine the most suitable sport. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions when using the system. The emotion engine analyzes and recognizes the user's emotions from inputs such as voice, facial expressions, and biosensors.
[0126] Based on information obtained from the emotion engine, the server adjusts the sports recommendation results considering the user's current emotional state. In this process, the server uses the emotion recognition results to customize the recommendations according to the user's preferences and reactions, generating an optimized sports recommendation list for each user.
[0127] For example, if the user's voice and facial expressions indicate a relaxed state, the server could broaden the range of sports suggested, offering a wider variety of options. Conversely, if tension is detected, the server could narrow down the suggestions or prioritize sports considered to be less mentally taxing.
[0128] Finally, the server-generated list of adaptive recommendations based on emotions, along with the reasons for their selection, is sent to the device. The device displays this information clearly to the user and provides feedback. In this way, a more personalized selection of sports, taking emotional considerations into account, becomes possible. Users can use these results to plan their child's sports activities and receive information to help their child enjoy sports more.
[0129] The following describes the processing flow.
[0130] Step 1:
[0131] The user uploads image data of the child's entire body to a dedicated application. The user pays attention to whether the child is relaxed when taking the picture, and also records their facial expressions and voice.
[0132] Step 2:
[0133] The terminal processes the received image data, performing pre-processing such as noise reduction and image size adjustment. This pre-processing prepares the image for analysis.
[0134] Step 3:
[0135] The device applies computer vision algorithms to extract skeletal features from images. It identifies the locations of the face, shoulders, elbows, knees, ankles, etc., and calculates the associated skeletal dimensions.
[0136] Step 4:
[0137] The terminal converts the extracted skeletal features into a standard format and sends them to the server. The data is encrypted before transmission to ensure security.
[0138] Step 5:
[0139] The server analyzes the user's voice and facial expressions simultaneously with the image data, and uses an emotion engine to recognize the user's emotional state. The recognized emotion data is then used for subsequent analysis.
[0140] Step 6:
[0141] The server compares the accumulated athlete data with the received skeletal features to determine the most suitable sport. This process utilizes machine learning algorithms to make the optimal selection based on skeletal attributes.
[0142] Step 7:
[0143] The server receives the results from the emotion engine and adjusts the recommended sports list according to the user's emotional state. For example, if the user is feeling stressed, sports with relaxing effects will be prioritized.
[0144] Step 8:
[0145] The server formats the adjusted sports recommendations and selection reasons for end users and sends them to their devices. The recommendations are customized to take into account the user's current mood.
[0146] Step 9:
[0147] The device decrypts the received data and displays it to the user in an easy-to-understand visual format. Based on this information, the user can make decisions about planning their child's sports activities.
[0148] (Example 2)
[0149] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0150] Currently, many systems suggest general exercises and activities to users, but these do not take into account the user's physical characteristics or emotional state. Therefore, they cannot suggest exercises that are optimal for individual needs, which can hinder user satisfaction and the ability to maintain effective activity. The challenge is to solve this problem and provide more personalized exercise recommendations.
[0151] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0152] In this invention, the server includes means for preprocessing image data, means for extracting skeletal features, and means for recognizing the user's emotional information. This enables personalized exercise recommendations based on the user's physical characteristics and emotional state.
[0153] "Image data" refers to data that represents visual information in a digital format, such as photographs and videos that are processed within a system.
[0154] "Preprocessing" refers to the initial stages of processing performed to make data easier to analyze, including tasks such as formatting standardization, noise reduction, and brightness adjustment.
[0155] "Skeletal features" refer to information about the positional relationships of joints and bones in the human body, extracted from image data.
[0156] "Adjustment" refers to the act of appropriately modifying data or results according to specific conditions or criteria, and in this context, it means optimizing exercise recommendations based on the user's emotional state.
[0157] The embodiment for carrying out the present invention is configured as a system that allows users to receive more personalized exercise suggestions.
[0158] Users run a dedicated application on their smartphone or tablet to acquire and upload image data of their child's entire body. This device has a function to preprocess the image data, performing image resizing, noise reduction, and brightness adjustment. This improves image quality and allows for accurate extraction of skeletal features.
[0159] The terminal extracts skeletal features from images using a computer vision algorithm. This algorithm, for example, uses an open-source computer vision library to identify joint points in a person. The extracted skeletal feature data is sent from the terminal to the server.
[0160] The server compares the received skeletal feature data with already stored human data to determine the optimal movement. A machine learning model runs on the server, rapidly optimizing based on the received data. Furthermore, the server integrates an emotion engine that analyzes the user's voice, facial expressions, and biometric data to extract emotional information. This analysis utilizes common speech recognition systems and facial expression analysis software.
[0161] The server adjusts the exercise recommendation list based on the user's emotional state. For example, it can broaden the exercise options if the user is relaxed, or prioritize specific exercises if they are stressed. This personalized exercise recommendation, along with the reasons for its selection, is ultimately sent to the user's device for visual confirmation.
[0162] For example, when the user is relaxed, the system is expected to recommend "extensive and challenging exercise." Conversely, if a state of tension is detected, it will recommend "easy and low-stress exercise." An example of a prompt to the generative AI model would be: "I would like to upload image data of my child and receive a sports recommendation. Please extract skeletal features and suggest the most suitable sport based on the user's emotional state. The user is currently relaxed."
[0163] This structure enables the present invention to recommend exercises that take into account the user's physical characteristics and emotional state, thereby supporting more effective and personalized health promotion.
[0164] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0165] Step 1:
[0166] The user uses a dedicated application to acquire image data of the child's entire body. The input is image data acquired from the camera, and the output is an image file saved on the device. This process prepares the basic data that the system will use for subsequent processing.
[0167] Step 2:
[0168] The device performs preprocessing on the acquired image data. The input is the image data acquired by the user, and the output is preprocessed, clean, and uniform image data. Specifically, the device performs image resizing, noise reduction, and brightness adjustment. This process allows for more effective and accurate extraction of skeletal features.
[0169] Step 3:
[0170] The device uses computer vision algorithms to extract skeletal features. The input is pre-processed image data, and the output is data on the joint positions of the identified skeleton. Specifically, it uses a library to detect points of the body within the image. The output obtained in this step becomes the data necessary for motion optimization.
[0171] Step 4:
[0172] The terminal sends extracted skeletal feature data to the server. The input is skeletal joint position data, and the output is data securely transferred to the server. A communication protocol is used to encode and compress the data, achieving efficient and secure data transfer.
[0173] Step 5:
[0174] The server compares the received skeletal feature data with a stored human database. The input is the transferred skeletal feature data, and the output is a movement suggestion based on the person's motor skills. A machine learning algorithm is used to analyze the data and determine which movement is appropriate.
[0175] Step 6:
[0176] The server uses an emotion engine to analyze the user's emotional information. Inputs include voice, facial expressions, and biometric data, while output is metadata indicating emotion. Emotion recognition software is used to identify emotions in real time, and the system adjusts accordingly.
[0177] Step 7:
[0178] The server adjusts exercise suggestions based on emotional information. Inputs are exercise suggestions and emotional metadata, while output is an exercise suggestion adapted to the user's emotions. An adjustment algorithm is executed to select an exercise that matches the user's current emotional state. This process enables the server to provide emotionally appropriate exercise guidance to the user.
[0179] Step 8:
[0180] The server sends the final exercise suggestions and the reasons for their selection to the terminal. The input is the adjusted list of exercise suggestions, and the output is the information displayed on the user's terminal. Data transfer technology is used to send the information to the terminal and display it in the most optimal format. As a result, the user receives specific exercise suggestions and can use them to implement them.
[0181] (Application Example 2)
[0182] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0183] In industrial environments, worker fatigue and emotional state can affect work efficiency and safety. Therefore, there is a need for methods to improve safety by monitoring workers' conditions in real time and providing work instructions tailored to their specific situations.
[0184] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0185] In this invention, the server includes means for acquiring image information, means for extracting skeletal features, and means for analyzing emotional states. This makes it possible to grasp the health status of workers in real time and provide appropriate work instructions.
[0186] "Image information" refers to information recorded in visual data format, including the images that are the subject of analysis.
[0187] "Skeletal features" refer to feature data extracted from image information that indicates the skeletal structure of a person or object.
[0188] "Athlete data" refers to a database of skeletal and related information of athletes accumulated over time, and is used as a standard for comparison.
[0189] A "user" is an individual or organization that obtains information or makes decisions through the system.
[0190] "Emotional state" refers to information that indicates an individual's current psychological and physiological state.
[0191] "Confidentiality" is the process of protecting information using specific technologies to prevent its misuse or leakage.
[0192] To realize this invention, a system is configured in which a server, terminal, and user work together. Specifically, the terminal is equipped with a high-resolution camera and microphone to acquire image information and audio data of the worker. The image information is preprocessed using an image processing library such as OpenCV to extract skeletal features. The extracted skeletal data is sent to the server and compared with previously accumulated data on athletes.
[0193] The server uses machine learning models built with PyTorch and TensorFlow to analyze the user's emotional state. Through data analysis, the server identifies the worker's fatigue level and stress level, and then derives optimal work instructions accordingly. If the server determines the user is relaxed, a wide range of tasks are suggested. Conversely, if tension is detected, calmer tasks are prioritized.
[0194] As a concrete example, factory workers are monitored by cameras and biosensors, and if fatigue is detected during work, the system prompts them to take appropriate breaks. It also recommends that workers switch from tasks requiring caution to safer tasks.
[0195] An example of a prompt for a generative AI model is: "Describe a program that uses video and audio data of a worker to analyze their skeletal structure and emotions, and then provides the worker with appropriate work instructions."
[0196] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0197] Step 1:
[0198] The terminal uses its built-in high-resolution camera and microphone to acquire image information and audio data of the worker. The input consists of camera images and audio data, which serve as the initial input data to the system. The output is sent for preprocessing in its original format.
[0199] Step 2:
[0200] The device performs preprocessing on the acquired image information using OpenCV, denoising and normalizing the image data. It also performs calculations to extract skeletal features from the image data. The input is raw image data, and the output is processed skeletal feature data.
[0201] Step 3:
[0202] The terminal sends skeletal feature data to the server. The server compares the received skeletal feature data with accumulated data on athletes. The input is skeletal feature data, and through comparison calculations, the optimal task list for the worker is generated. The output is the optimal task list.
[0203] Step 4:
[0204] The server uses PyTorch or TensorFlow to perform emotion analysis using voice and other sensor data. Inputs are voice data and, if necessary, sensor data. Through this analysis, the server determines and outputs the worker's emotional state.
[0205] Step 5:
[0206] The server integrates skeletal feature comparison results and sentiment analysis results to customize tasks to suit the worker. The input is the optimal task list and sentiment analysis results, and adjusted work instructions are generated based on this. The output is the adjusted work instruction list.
[0207] Step 6:
[0208] The adjusted work order list is sent to the terminal and displayed to the user. Based on this information, the user can perform or modify the work. The input is the adjusted work order list, and the output is the work based on the user's decisions.
[0209] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0210] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0211] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0212] [Second Embodiment]
[0213] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0214] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0215] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0216] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0217] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0218] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0219] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0220] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0221] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0222] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0223] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0224] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0225] To implement this invention, a user acquires image data of a child's entire body using a dedicated application. Upon receiving this image data, the terminal first performs preprocessing, such as noise reduction and size adjustment. Subsequently, it extracts skeletal features from the image using computer vision technology. Specifically, it identifies the positions of the face, shoulders, elbows, knees, and ankles, and calculates their relative distances and angles.
[0226] The extracted skeletal features are sent to the server in a standardized format. Upon receiving these skeletal features, the server compares them to a database of athletes and applies a machine learning algorithm to determine the most suitable sport. The server refers to the profiles of numerous athletes stored in the database and identifies sports with statistically similar skeletal characteristics. Based on this analysis, it generates a list of the most suitable sports and the reasons for their selection, and sends it to the user's terminal.
[0227] For example, suppose an analysis of image data of a specific child reveals that the child is tall and has long arms. The server determines that these characteristics are highly correlated with sports where they are typically advantageous, such as basketball, volleyball, and certain track and field events. It then generates a recommendation list that includes these sports and sends it to the terminal with an explanation such as "height and arm length are related."
[0228] In this way, the system of the present invention provides a means to enable users to make scientifically and objectively rational choices.
[0229] Based on the information received, users can assess their child's aptitude and create a concrete plan for sports activities. This helps unlock the child's potential and provides them with opportunities to enjoy sports in an optimal environment.
[0230] The following describes the processing flow.
[0231] Step 1:
[0232] The user uploads image data of the child's entire body to a dedicated application. The user ensures that the child is sitting upright when the photo is taken.
[0233] Step 2:
[0234] The device preprocesses the image data it receives. Specifically, it adjusts the image resolution and removes noise as needed to prepare the image for analysis.
[0235] Step 3:
[0236] The device uses a computer vision algorithm to extract skeletal features from image data. The algorithm identifies the location of each point in the skeleton, such as the shoulders, elbows, knees, and ankles, and calculates the distances and angles between these locations.
[0237] Step 4:
[0238] The terminal converts the extracted skeletal features into a standardized format, encrypts them for security, and then sends them to the server.
[0239] Step 5:
[0240] The server analyzes the skeletal features it receives and compares them to accumulated athlete data. Here, a machine learning algorithm is used to identify the sport with the highest correlation.
[0241] Step 6:
[0242] Based on the analysis results, the server generates a list of the most suitable sports for the child, along with the reasons for their selection. This includes explaining why certain skeletal characteristics are advantageous for specific sports.
[0243] Step 7:
[0244] The server encrypts the generated sports recommendation list and sends it to the terminal. Privacy protection is taken into consideration during this process.
[0245] Step 8:
[0246] The device decrypts the received data and displays it in an easy-to-understand format for the user. Based on this, the user can then plan their child's sports activities.
[0247] (Example 1)
[0248] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0249] Conventional sports aptitude assessment systems suffered from insufficient preprocessing, data transmission, and analysis methods in the process from image data acquisition to aptitude assessment, resulting in challenges in terms of prediction accuracy and security. Furthermore, the feedback provided to users before suggesting appropriate exercise activities was inadequate, sometimes failing to deliver reliable judgments.
[0250] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0251] In this invention, the server includes means for preprocessing image information, means for transmitting joint features in a standard format, and means for analyzing the data using a clustering algorithm. This enables effective and reliable assessment of suitability for exercise activities by statistically comparing it with data from athletes.
[0252] "Image information" refers to visual data acquired in digital format that includes the visual characteristics of a specific object.
[0253] "Preprocessing" refers to the process of applying noise reduction, size adjustment, and other adjustments to acquired image information to prepare the data for a useful format.
[0254] "Joint features" refer to data extracted from image information that shows the location of joints in the human body and their spatial relationships.
[0255] A "standard format" is a structured data format that is consistently used in the transmission, reception, and processing of data.
[0256] A "clustering algorithm" refers to a computational method used to statistically analyze data and group elements that have similar characteristics.
[0257] "Athlete data" refers to a collection of information regarding the physical characteristics and performance of athletes engaged in diverse athletic activities.
[0258] "Assessing suitability for physical activity" is the act of identifying the most suitable exercise or sport for an individual based on analyzed data.
[0259] To implement this invention, the user first uses a dedicated application to acquire full-body image information of a child using a camera-equipped device. The terminal then performs preprocessing on the received image information, such as noise reduction and size adjustment, using an image processing library. Specifically, the terminal utilizes image processing software such as OpenCV to optimize the image quality.
[0260] Next, the device utilizes computer vision technology to extract joint features from the image. This process uses OpenPose, a pre-trained model, to calculate the coordinates of the face and joint positions. This information is structured in a standard format such as JSON and treated as data.
[0261] The terminal sends standardized joint feature data to the server using the HTTPS protocol. After receiving this data, the server applies a clustering algorithm to compare it with the accumulated data on athletes. The server uses machine learning libraries such as Scikit-learn and TensorFlow, employing techniques such as K-means clustering, to identify sports with features similar to the input data.
[0262] Based on this analysis, the server generates a list of suitable sports and the reasons for their selection, and sends the information back to the user's device. The user can then review these results on their device and plan exercise activities that are tailored to their child's characteristics.
[0263] For example, if an analysis of a child's image data reveals features such as being tall and having long arms, the server will determine that these features are advantageous in sports like basketball or volleyball. This determination is then returned to the user with a reason for the selection, such as "because height and arm length are correlated."
[0264] An example of a prompt statement can be expressed as follows:
[0265] "Please provide information on a system that suggests optimal exercise activities based on joint features identified from images of children. Specifically, I'd like to know which features correlate with which sports, and how the server analyzes and returns the results."
[0266] Thus, the system of the present invention provides a means to support sports selection in a scientific and objective manner.
[0267] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0268] Step 1:
[0269] The user uses a dedicated application to capture full-body images of their child using the camera on their smartphone or tablet. The images are saved in JPEG or PNG format. The captured image information is then transmitted to the device through the application's interface.
[0270] Step 2:
[0271] The terminal performs preprocessing on the received image information. It receives an image in JPEG or PNG format as input, uses an image processing library to remove noise (using a Gaussian filter) and adjust the size (resizing to 224x224 pixels), and generates a clean, appropriately sized image as output.
[0272] Step 3:
[0273] The device extracts joint features using pre-processed images. Using pre-processed images as input, it detects the positions of faces and joints using computer vision techniques such as OpenPose. As output, it generates a dataset containing coordinate information for each joint.
[0274] Step 4:
[0275] The terminal converts the extracted joint features into a standard format (such as JSON). It uses joint coordinate information as input and converts this data into a structured format such as JSON. As output, it generates a data package that can be sent to the server.
[0276] Step 5:
[0277] The terminal sends a data package of joint features to the server. Using the HTTPS protocol, the data is securely transferred to the server, and once the transmission is complete, the terminal enters a state of waiting for a response from the server.
[0278] Step 6:
[0279] The server performs analysis based on the received joint feature data. Using the received coordinate data as input, it analyzes the data using algorithms such as K-means clustering, utilizing machine learning libraries (Scikit-learn, TensorFlow, etc.). As output, it generates a list of exercise activities suitable for the subject.
[0280] Step 7:
[0281] The server generates analysis results (a list of suitable exercise activities and the reasons for them), encrypts them, and sends them to the terminal. The output provides data in a format viewable on the user's terminal.
[0282] Step 8:
[0283] The user receives feedback from the server on the terminal and checks the presented list of suitable sports. Based on this information, the user can plan the child's sports activities.
[0284] (Application Example 1)
[0285] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0286] In the delivery service, the problem is that the delivery staff cannot select the most efficient delivery area or route based on their individual physical characteristics. As a result, their individual abilities cannot be fully utilized, and there is a risk of a decrease in work efficiency. The purpose of the present invention is to eliminate such inefficiencies and provide a system that proposes an appropriate delivery process for the delivery staff.
[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0288] In this invention, the server includes means for acquiring image data, means for extracting physical characteristics from the image data, means for comparing the physical characteristics with individual data accumulated in advance and determining an optimal behavior pattern, means for presenting the determination result to the user, and means for recommending an efficient area or route based on the determination result. As a result, it becomes possible to propose an optimal delivery area or route according to the delivery staff.
[0289] <00A "behavioral pattern" is a pattern of optimal behavior or activity that is predicted based on an individual's characteristics.
[0293] "Users" refer to entities or individuals that utilize this system.
[0294] "Region or route" refers to the specific locations or travel routes where the delivery person performs their duties.
[0295] "Recommendation" means presenting the option that best suits specific conditions.
[0296] In implementing this invention, the user first uses a terminal such as a smartphone to acquire whole-body image data of the target individual. The terminal utilizes advanced image processing techniques to extract physical characteristics from the image data. For this processing, OpenCV, an open-source computer vision library, is used for noise reduction and size adjustment.
[0297] The physical characteristics data extracted by the device is sent to the server. The server compares the received physical characteristics with existing data in the individual database and uses a generative AI model to determine the optimal behavioral pattern. During this process, the server performs data analysis using machine learning libraries such as TensorFlow and PyTorch.
[0298] The results are presented to the user. These results include the most suitable region or route for the individual, and suggest efficient actions and activities.
[0299] For example, if a delivery person has physical characteristics such as being tall and having long arms, the server will determine that they are suitable for long-distance deliveries and suggest a delivery route that covers a wide area. This suggestion is then notified to the user's smartphone.
[0300] As an example of a prompt sentence when inputting into the generation AI model, a format such as "Based on the physical characteristics of the delivery person, please generate an efficient delivery route. For example, for a delivery person who is tall and has long arms, please propose a route that covers a wide area" can be considered. In this way, the present invention provides a system that enables an efficient operation suitable for the characteristics of the delivery person.
[0301] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0302] Step 1:
[0303] The user uses the camera of the smartphone to take a full-body image of the delivery person. The input is the captured full-body image data, and the output is the image file saved in the terminal. The terminal starts processing with this image data as the initial input.
[0304] Step 2:
[0305] The terminal uses the OpenCV library to perform preprocessing on the acquired image data. The input is the full-body image data, and filtering and resizing processes are performed to remove noise from the image and adjust the size. The output is the preprocessed clear image data. Specifically, the sharpness of the image is improved and the resolution is adjusted for easy analysis.
[0306] Step 3:
[0307] <00性情>000968>Based on the preprocessed image data, the terminal uses TensorFlow to extract physical characteristics. The input is the preprocessed image data, and the output is physical characteristic data such as the position and length of the skeleton. For this data extraction, a pre-trained model is used to identify the positions of the face, shoulders, elbows, knees, and ankles, and calculate their relative distances and angles.
[0308] Step 4: [[ID=The terminal sends extracted physical characteristic data to the server. The server receives this data and compares it with existing data in the individual database. The input is physical characteristic data, and the output is identification information for similar database entries. Using a database search algorithm, the server applies a statistical model to efficiently determine similarity.
[0310] Step 5:
[0311] The server runs a machine learning model and compares it with an individual database to determine the optimal behavioral pattern. The input is the physical characteristics data received by the server and the information in the database, and the output is a suggestion of the optimal region or route specific to the individual. This model uses a deep learning algorithm implemented in TensorFlow or PyTorch.
[0312] Step 6:
[0313] The server sends the determined suggestion to the user. The input is the region or route information determined by the server, and the output is a notification message displayed on the user's device. The notification includes a brief explanation of the delivery route details and efficiency. Specifically, the notification message appears as a pop-up on the user's smartphone.
[0314] Step 7:
[0315] Users review notifications and perform delivery tasks according to the suggested area or route. The input is the information presented as a notification, and the output is the streamlined actual delivery task. By adhering to these instructions, users can expect improved operational efficiency and minimized effort.
[0316] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0317] This invention begins with a user using a dedicated application to acquire image data of a child's entire body and uploading this data to a terminal. The image data is preprocessed on the terminal and processed by a computer vision algorithm to extract skeletal features. After detecting the skeletal features, the terminal sends the feature data to a server.
[0318] The server uses the received skeletal features to compare them with pre-stored athlete data and determine the most suitable sport. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions when using the system. The emotion engine analyzes and recognizes the user's emotions from inputs such as voice, facial expressions, and biosensors.
[0319] Based on information obtained from the emotion engine, the server adjusts the sports recommendation results considering the user's current emotional state. In this process, the server uses the emotion recognition results to customize the recommendations according to the user's preferences and reactions, generating an optimized sports recommendation list for each user.
[0320] For example, if the user's voice and facial expressions indicate a relaxed state, the server could broaden the range of sports suggested, offering a wider variety of options. Conversely, if tension is detected, the server could narrow down the suggestions or prioritize sports considered to be less mentally taxing.
[0321] Finally, the server-generated list of adaptive recommendations based on emotions, along with the reasons for their selection, is sent to the device. The device displays this information clearly to the user and provides feedback. In this way, a more personalized selection of sports, taking emotional considerations into account, becomes possible. Users can use these results to plan their child's sports activities and receive information to help their child enjoy sports more.
[0322] The following describes the processing flow.
[0323] Step 1:
[0324] The user uploads image data of the child's entire body to a dedicated application. The user pays attention to whether the child is relaxed when taking the picture, and also records their facial expressions and voice.
[0325] Step 2:
[0326] The terminal processes the received image data, performing pre-processing such as noise reduction and image size adjustment. This pre-processing prepares the image for analysis.
[0327] Step 3:
[0328] The device applies computer vision algorithms to extract skeletal features from images. It identifies the locations of the face, shoulders, elbows, knees, ankles, etc., and calculates the associated skeletal dimensions.
[0329] Step 4:
[0330] The terminal converts the extracted skeletal features into a standard format and sends them to the server. The data is encrypted before transmission to ensure security.
[0331] Step 5:
[0332] The server analyzes the user's voice and facial expressions simultaneously with the image data, and uses an emotion engine to recognize the user's emotional state. The recognized emotion data is then used for subsequent analysis.
[0333] Step 6:
[0334] The server compares the accumulated athlete data with the received skeletal features to determine the most suitable sport. This process utilizes machine learning algorithms to make the optimal selection based on skeletal attributes.
[0335] Step 7:
[0336] The server receives the results from the emotion engine and adjusts the recommended sports list according to the user's emotional state. For example, if the user is feeling stressed, sports with relaxing effects will be prioritized.
[0337] Step 8:
[0338] The server formats the adjusted sports recommendations and selection reasons for end users and sends them to their devices. The recommendations are customized to take into account the user's current mood.
[0339] Step 9:
[0340] The device decrypts the received data and displays it to the user in an easy-to-understand visual format. Based on this information, the user can make decisions about planning their child's sports activities.
[0341] (Example 2)
[0342] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0343] Currently, many systems suggest general exercises and activities to users, but these do not take into account the user's physical characteristics or emotional state. Therefore, they cannot suggest exercises that are optimal for individual needs, which can hinder user satisfaction and the ability to maintain effective activity. The challenge is to solve this problem and provide more personalized exercise recommendations.
[0344] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0345] In this invention, the server includes means for preprocessing image data, means for extracting skeletal features, and means for recognizing the user's emotional information. This enables personalized exercise recommendations based on the user's physical characteristics and emotional state.
[0346] "Image data" refers to data that represents visual information in a digital format, such as photographs and videos that are processed within a system.
[0347] "Preprocessing" refers to the initial stages of processing performed to make data easier to analyze, including tasks such as formatting standardization, noise reduction, and brightness adjustment.
[0348] "Skeletal features" refer to information about the positional relationships of joints and bones in the human body, extracted from image data.
[0349] "Adjustment" refers to the act of appropriately modifying data or results according to specific conditions or criteria, and in this context, it means optimizing exercise recommendations based on the user's emotional state.
[0350] The embodiment for carrying out the present invention is configured as a system that allows users to receive more personalized exercise suggestions.
[0351] Users run a dedicated application on their smartphone or tablet to acquire and upload image data of their child's entire body. This device has a function to preprocess the image data, performing image resizing, noise reduction, and brightness adjustment. This improves image quality and allows for accurate extraction of skeletal features.
[0352] The terminal extracts skeletal features from images using a computer vision algorithm. This algorithm, for example, uses an open-source computer vision library to identify joint points in a person. The extracted skeletal feature data is sent from the terminal to the server.
[0353] The server compares the received skeletal feature data with already stored human data to determine the optimal movement. A machine learning model runs on the server, rapidly optimizing based on the received data. Furthermore, the server integrates an emotion engine that analyzes the user's voice, facial expressions, and biometric data to extract emotional information. This analysis utilizes common speech recognition systems and facial expression analysis software.
[0354] The server adjusts the exercise recommendation list based on the user's emotional state. For example, it can broaden the exercise options if the user is relaxed, or prioritize specific exercises if they are stressed. This personalized exercise recommendation, along with the reasons for its selection, is ultimately sent to the user's device for visual confirmation.
[0355] For example, when the user is relaxed, the system is expected to recommend "extensive and challenging exercise." Conversely, if a state of tension is detected, it will recommend "easy and low-stress exercise." An example of a prompt to the generative AI model would be: "I would like to upload image data of my child and receive a sports recommendation. Please extract skeletal features and suggest the most suitable sport based on the user's emotional state. The user is currently relaxed."
[0356] This structure enables the present invention to recommend exercises that take into account the user's physical characteristics and emotional state, thereby supporting more effective and personalized health promotion.
[0357] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0358] Step 1:
[0359] The user uses a dedicated application to acquire image data of the child's entire body. The input is image data acquired from the camera, and the output is an image file saved on the device. This process prepares the basic data that the system will use for subsequent processing.
[0360] Step 2:
[0361] The device performs preprocessing on the acquired image data. The input is the image data acquired by the user, and the output is preprocessed, clean, and uniform image data. Specifically, the device performs image resizing, noise reduction, and brightness adjustment. This process allows for more effective and accurate extraction of skeletal features.
[0362] Step 3:
[0363] The device uses computer vision algorithms to extract skeletal features. The input is pre-processed image data, and the output is data on the joint positions of the identified skeleton. Specifically, it uses a library to detect points of the body within the image. The output obtained in this step becomes the data necessary for motion optimization.
[0364] Step 4:
[0365] The terminal sends extracted skeletal feature data to the server. The input is skeletal joint position data, and the output is data securely transferred to the server. A communication protocol is used to encode and compress the data, achieving efficient and secure data transfer.
[0366] Step 5:
[0367] The server compares the received skeletal feature data with a stored human database. The input is the transferred skeletal feature data, and the output is a movement suggestion based on the person's motor skills. A machine learning algorithm is used to analyze the data and determine which movement is appropriate.
[0368] Step 6:
[0369] The server uses an emotion engine to analyze the user's emotional information. Inputs include voice, facial expressions, and biometric data, while output is metadata indicating emotion. Emotion recognition software is used to identify emotions in real time, and the system adjusts accordingly.
[0370] Step 7:
[0371] The server adjusts exercise suggestions based on emotional information. Inputs are exercise suggestions and emotional metadata, while output is an exercise suggestion adapted to the user's emotions. An adjustment algorithm is executed to select an exercise that matches the user's current emotional state. This process enables the server to provide emotionally appropriate exercise guidance to the user.
[0372] Step 8:
[0373] The server sends the final exercise suggestions and the reasons for their selection to the terminal. The input is the adjusted list of exercise suggestions, and the output is the information displayed on the user's terminal. Data transfer technology is used to send the information to the terminal and display it in the most optimal format. As a result, the user receives specific exercise suggestions and can use them to implement them.
[0374] (Application Example 2)
[0375] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0376] In industrial environments, worker fatigue and emotional state can affect work efficiency and safety. Therefore, there is a need for methods to improve safety by monitoring workers' conditions in real time and providing work instructions tailored to their specific situations.
[0377] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0378] In this invention, the server includes means for acquiring image information, means for extracting skeletal features, and means for analyzing emotional states. This makes it possible to grasp the health status of workers in real time and provide appropriate work instructions.
[0379] "Image information" refers to information recorded in visual data format, including the images that are the subject of analysis.
[0380] "Skeletal features" refer to feature data extracted from image information that indicates the skeletal structure of a person or object.
[0381] "Athlete data" refers to a database of skeletal and related information of athletes accumulated over time, and is used as a standard for comparison.
[0382] A "user" is an individual or organization that obtains information or makes decisions through the system.
[0383] "Emotional state" refers to information that indicates an individual's current psychological and physiological state.
[0384] "Confidentiality" is the process of protecting information using specific technologies to prevent its misuse or leakage.
[0385] To realize this invention, a system is configured in which a server, terminal, and user work together. Specifically, the terminal is equipped with a high-resolution camera and microphone to acquire image information and audio data of the worker. The image information is preprocessed using an image processing library such as OpenCV to extract skeletal features. The extracted skeletal data is sent to the server and compared with previously accumulated data on athletes.
[0386] The server uses machine learning models built with PyTorch and TensorFlow to analyze the user's emotional state. Through data analysis, the server identifies the worker's fatigue level and stress level, and then derives optimal work instructions accordingly. If the server determines the user is relaxed, a wide range of tasks are suggested. Conversely, if tension is detected, calmer tasks are prioritized.
[0387] As a concrete example, factory workers are monitored by cameras and biosensors, and if fatigue is detected during work, the system prompts them to take appropriate breaks. It also recommends that workers switch from tasks requiring caution to safer tasks.
[0388] An example of a prompt for a generative AI model is: "Describe a program that uses video and audio data of a worker to analyze their skeletal structure and emotions, and then provides the worker with appropriate work instructions."
[0389] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0390] Step 1:
[0391] The terminal uses its built-in high-resolution camera and microphone to acquire image information and audio data of the worker. The input consists of camera images and audio data, which serve as the initial input data to the system. The output is sent for preprocessing in its original format.
[0392] Step 2:
[0393] The device performs preprocessing on the acquired image information using OpenCV, denoising and normalizing the image data. It also performs calculations to extract skeletal features from the image data. The input is raw image data, and the output is processed skeletal feature data.
[0394] Step 3:
[0395] The terminal sends skeletal feature data to the server. The server compares the received skeletal feature data with accumulated data on athletes. The input is skeletal feature data, and through comparison calculations, the optimal task list for the worker is generated. The output is the optimal task list.
[0396] Step 4:
[0397] The server uses PyTorch or TensorFlow to perform emotion analysis using voice and other sensor data. Inputs are voice data and, if necessary, sensor data. Through this analysis, the server determines and outputs the worker's emotional state.
[0398] Step 5:
[0399] The server integrates skeletal feature comparison results and sentiment analysis results to customize tasks to suit the worker. The input is the optimal task list and sentiment analysis results, and adjusted work instructions are generated based on this. The output is the adjusted work instruction list.
[0400] Step 6:
[0401] The adjusted work order list is sent to the terminal and displayed to the user. Based on this information, the user can perform or modify the work. The input is the adjusted work order list, and the output is the work based on the user's decisions.
[0402] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0403] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0404] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0405] [Third Embodiment]
[0406] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0407] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0408] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0409] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0410] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0411] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0412] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0413] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0414] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0415] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0416] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0417] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0418] To implement this invention, a user acquires image data of a child's entire body using a dedicated application. Upon receiving this image data, the terminal first performs preprocessing, such as noise reduction and size adjustment. Subsequently, it extracts skeletal features from the image using computer vision technology. Specifically, it identifies the positions of the face, shoulders, elbows, knees, and ankles, and calculates their relative distances and angles.
[0419] The extracted skeletal features are sent to the server in a standardized format. Upon receiving these skeletal features, the server compares them to a database of athletes and applies a machine learning algorithm to determine the most suitable sport. The server refers to the profiles of numerous athletes stored in the database and identifies sports with statistically similar skeletal characteristics. Based on this analysis, it generates a list of the most suitable sports and the reasons for their selection, and sends it to the user's terminal.
[0420] For example, suppose an analysis of image data of a specific child reveals that the child is tall and has long arms. The server determines that these characteristics are highly correlated with sports where they are typically advantageous, such as basketball, volleyball, and certain track and field events. It then generates a recommendation list that includes these sports and sends it to the terminal with an explanation such as "height and arm length are related."
[0421] In this way, the system of the present invention provides a means to enable users to make scientifically and objectively rational choices.
[0422] Based on the information received, users can assess their child's aptitude and create a concrete plan for sports activities. This helps unlock the child's potential and provides them with opportunities to enjoy sports in an optimal environment.
[0423] The following describes the processing flow.
[0424] Step 1:
[0425] The user uploads image data of the child's entire body to a dedicated application. The user ensures that the child is sitting upright when the photo is taken.
[0426] Step 2:
[0427] The device preprocesses the image data it receives. Specifically, it adjusts the image resolution and removes noise as needed to prepare the image for analysis.
[0428] Step 3:
[0429] The device uses a computer vision algorithm to extract skeletal features from image data. The algorithm identifies the location of each point in the skeleton, such as the shoulders, elbows, knees, and ankles, and calculates the distances and angles between these locations.
[0430] Step 4:
[0431] The terminal converts the extracted skeletal features into a standardized format, encrypts them for security, and then sends them to the server.
[0432] Step 5:
[0433] The server analyzes the skeletal features it receives and compares them to accumulated athlete data. Here, a machine learning algorithm is used to identify the sport with the highest correlation.
[0434] Step 6:
[0435] Based on the analysis results, the server generates a list of the most suitable sports for the child, along with the reasons for their selection. This includes explaining why certain skeletal characteristics are advantageous for specific sports.
[0436] Step 7:
[0437] The server encrypts the generated sports recommendation list and sends it to the terminal. Privacy protection is taken into consideration during this process.
[0438] Step 8:
[0439] The device decrypts the received data and displays it in an easy-to-understand format for the user. Based on this, the user can then plan their child's sports activities.
[0440] (Example 1)
[0441] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0442] Conventional sports aptitude assessment systems suffered from insufficient preprocessing, data transmission, and analysis methods in the process from image data acquisition to aptitude assessment, resulting in challenges in terms of prediction accuracy and security. Furthermore, the feedback provided to users before suggesting appropriate exercise activities was inadequate, sometimes failing to deliver reliable judgments.
[0443] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0444] In this invention, the server includes means for preprocessing image information, means for transmitting joint features in a standard format, and means for analyzing the data using a clustering algorithm. This enables effective and reliable assessment of suitability for exercise activities by statistically comparing it with data from athletes.
[0445] "Image information" refers to visual data acquired in digital format that includes the visual characteristics of a specific object.
[0446] "Preprocessing" refers to the process of applying noise reduction, size adjustment, and other adjustments to acquired image information to prepare the data for a useful format.
[0447] "Joint features" refer to data extracted from image information that shows the location of joints in the human body and their spatial relationships.
[0448] A "standard format" is a structured data format that is consistently used in the transmission, reception, and processing of data.
[0449] A "clustering algorithm" refers to a computational method used to statistically analyze data and group elements that have similar characteristics.
[0450] "Athlete data" refers to a collection of information regarding the physical characteristics and performance of athletes engaged in diverse athletic activities.
[0451] "Assessing suitability for physical activity" is the act of identifying the most suitable exercise or sport for an individual based on analyzed data.
[0452] To implement this invention, the user first uses a dedicated application to acquire full-body image information of a child using a camera-equipped device. The terminal then performs preprocessing on the received image information, such as noise reduction and size adjustment, using an image processing library. Specifically, the terminal utilizes image processing software such as OpenCV to optimize the image quality.
[0453] Next, the device utilizes computer vision technology to extract joint features from the image. This process uses OpenPose, a pre-trained model, to calculate the coordinates of the face and joint positions. This information is structured in a standard format such as JSON and treated as data.
[0454] The terminal sends standardized joint feature data to the server using the HTTPS protocol. After receiving this data, the server applies a clustering algorithm to compare it with the accumulated data on athletes. The server uses machine learning libraries such as Scikit-learn and TensorFlow, employing techniques such as K-means clustering, to identify sports with features similar to the input data.
[0455] Based on this analysis, the server generates a list of suitable sports and the reasons for their selection, and sends the information back to the user's device. The user can then review these results on their device and plan exercise activities that are tailored to their child's characteristics.
[0456] For example, if an analysis of a child's image data reveals features such as being tall and having long arms, the server will determine that these features are advantageous in sports like basketball or volleyball. This determination is then returned to the user with a reason for the selection, such as "because height and arm length are correlated."
[0457] An example of a prompt statement can be expressed as follows:
[0458] "Please provide information on a system that suggests optimal exercise activities based on joint features identified from images of children. Specifically, I'd like to know which features correlate with which sports, and how the server analyzes and returns the results."
[0459] Thus, the system of the present invention provides a means to support sports selection in a scientific and objective manner.
[0460] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0461] Step 1:
[0462] The user uses a dedicated application to capture full-body images of their child using the camera on their smartphone or tablet. The images are saved in JPEG or PNG format. The captured image information is then transmitted to the device through the application's interface.
[0463] Step 2:
[0464] The terminal performs preprocessing on the received image information. It receives an image in JPEG or PNG format as input, uses an image processing library to remove noise (using a Gaussian filter) and adjust the size (resizing to 224x224 pixels), and generates a clean, appropriately sized image as output.
[0465] Step 3:
[0466] The device extracts joint features using pre-processed images. Using pre-processed images as input, it detects the positions of faces and joints using computer vision techniques such as OpenPose. As output, it generates a dataset containing coordinate information for each joint.
[0467] Step 4:
[0468] The terminal converts the extracted joint features into a standard format (such as JSON). It uses joint coordinate information as input and converts this data into a structured format such as JSON. As output, it generates a data package that can be sent to the server.
[0469] Step 5:
[0470] The terminal sends a data package of joint features to the server. Using the HTTPS protocol, the data is securely transferred to the server, and once the transmission is complete, the terminal enters a state of waiting for a response from the server.
[0471] Step 6:
[0472] The server performs analysis based on the received joint feature data. Using the received coordinate data as input, it analyzes the data using algorithms such as K-means clustering, utilizing machine learning libraries (Scikit-learn, TensorFlow, etc.). As output, it generates a list of exercise activities suitable for the subject.
[0473] Step 7:
[0474] The server generates analysis results (a list of suitable exercise activities and the reasons for them), encrypts them, and sends them to the terminal. The output provides data in a format viewable on the user's terminal.
[0475] Step 8:
[0476] Users receive feedback from the server on their device and review a list of suitable sports. Based on this information, they can plan their child's sports activities.
[0477] (Application Example 1)
[0478] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0479] In delivery services, a challenge is that delivery personnel are unable to select the most efficient delivery areas and routes based on their individual physical characteristics. As a result, they may not be able to make the most of their individual abilities, potentially leading to decreased work efficiency. This invention aims to eliminate this inefficiency and provide a system that proposes an appropriate delivery process for delivery personnel.
[0480] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0481] In this invention, the server includes means for acquiring image data, means for extracting physical characteristics from the image data, means for comparing the physical characteristics with pre-stored individual data to determine the optimal behavioral pattern, means for presenting the determination result to the user, and means for recommending an efficient region or route based on the determination result. This makes it possible to suggest the optimal delivery region and route for each delivery person.
[0482] "Image data" refers to a record of the visual information of an object in numerical format.
[0483] "Physical characteristics" refer to data that describes an individual's skeletal structure and physical features.
[0484] "Individual data" refers to a collection of information that accumulates the physical characteristics of multiple individuals.
[0485] A "behavioral pattern" is a pattern of optimal behavior or activity that is predicted based on an individual's characteristics.
[0486] "Users" refer to entities or individuals that utilize this system.
[0487] "Region or route" refers to the specific locations or travel routes where the delivery person performs their duties.
[0488] "Recommendation" means presenting the option that best suits specific conditions.
[0489] In implementing this invention, the user first uses a terminal such as a smartphone to acquire whole-body image data of the target individual. The terminal utilizes advanced image processing techniques to extract physical characteristics from the image data. For this processing, OpenCV, an open-source computer vision library, is used for noise reduction and size adjustment.
[0490] The physical characteristics data extracted by the device is sent to the server. The server compares the received physical characteristics with existing data in the individual database and uses a generative AI model to determine the optimal behavioral pattern. During this process, the server performs data analysis using machine learning libraries such as TensorFlow and PyTorch.
[0491] The results are presented to the user. These results include the most suitable region or route for the individual, and suggest efficient actions and activities.
[0492] For example, if a delivery person has physical characteristics such as being tall and having long arms, the server will determine that they are suitable for long-distance deliveries and suggest a delivery route that covers a wide area. This suggestion is then notified to the user's smartphone.
[0493] An example of a prompt message to be input into the generating AI model is: "Generate an efficient delivery route based on the physical characteristics of the delivery person. For example, for a tall delivery person with long arms, suggest a route that covers a wide area." In this way, the present invention provides a system that enables efficient work suited to the characteristics of the delivery person.
[0494] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0495] Step 1:
[0496] The user uses their smartphone camera to take a full-body image of the delivery person. The input is the captured full-body image data, and the output is an image file stored on the device. The device starts processing this image data as its initial input.
[0497] Step 2:
[0498] The terminal uses the OpenCV library to preprocess the acquired image data. The input is whole-body image data, and filtering and resizing are performed to remove noise and adjust the size of the image. The output is preprocessed, clear image data. Specifically, the image sharpness is improved and the resolution is adjusted to facilitate analysis.
[0499] Step 3:
[0500] The device uses TensorFlow to extract body characteristics based on pre-processed image data. The input is pre-processed image data, and the output is body characteristic data such as the position and length of the skeleton. This data extraction uses a pre-trained model to identify the positions of the face, shoulders, elbows, knees, and ankles, and to calculate their relative distances and angles.
[0501] Step 4:
[0502] The terminal sends extracted physical characteristic data to the server. The server receives this data and compares it with existing data in the individual database. The input is physical characteristic data, and the output is identification information for similar database entries. Using a database search algorithm, the server applies a statistical model to efficiently determine similarity.
[0503] Step 5:
[0504] The server runs a machine learning model and compares it with an individual database to determine the optimal behavioral pattern. The input is the physical characteristics data received by the server and the information in the database, and the output is a suggestion of the optimal region or route specific to the individual. This model uses a deep learning algorithm implemented in TensorFlow or PyTorch.
[0505] Step 6:
[0506] The server sends the determined suggestion to the user. The input is the region or route information determined by the server, and the output is a notification message displayed on the user's device. The notification includes a brief explanation of the delivery route details and efficiency. Specifically, the notification message appears as a pop-up on the user's smartphone.
[0507] Step 7:
[0508] Users review notifications and perform delivery tasks according to the suggested area or route. The input is the information presented as a notification, and the output is the streamlined actual delivery task. By adhering to these instructions, users can expect improved operational efficiency and minimized effort.
[0509] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0510] This invention begins with a user using a dedicated application to acquire image data of a child's entire body and uploading this data to a terminal. The image data is preprocessed on the terminal and processed by a computer vision algorithm to extract skeletal features. After detecting the skeletal features, the terminal sends the feature data to a server.
[0511] The server uses the received skeletal features to compare them with pre-stored athlete data and determine the most suitable sport. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions when using the system. The emotion engine analyzes and recognizes the user's emotions from inputs such as voice, facial expressions, and biosensors.
[0512] Based on information obtained from the emotion engine, the server adjusts the sports recommendation results considering the user's current emotional state. In this process, the server uses the emotion recognition results to customize the recommendations according to the user's preferences and reactions, generating an optimized sports recommendation list for each user.
[0513] For example, if the user's voice and facial expressions indicate a relaxed state, the server could broaden the range of sports suggested, offering a wider variety of options. Conversely, if tension is detected, the server could narrow down the suggestions or prioritize sports considered to be less mentally taxing.
[0514] Finally, the server-generated list of adaptive recommendations based on emotions, along with the reasons for their selection, is sent to the device. The device displays this information clearly to the user and provides feedback. In this way, a more personalized selection of sports, taking emotional considerations into account, becomes possible. Users can use these results to plan their child's sports activities and receive information to help their child enjoy sports more.
[0515] The following describes the processing flow.
[0516] Step 1:
[0517] The user uploads image data of the child's entire body to a dedicated application. The user pays attention to whether the child is relaxed when taking the picture, and also records their facial expressions and voice.
[0518] Step 2:
[0519] The terminal processes the received image data, performing pre-processing such as noise reduction and image size adjustment. This pre-processing prepares the image for analysis.
[0520] Step 3:
[0521] The device applies computer vision algorithms to extract skeletal features from images. It identifies the locations of the face, shoulders, elbows, knees, ankles, etc., and calculates the associated skeletal dimensions.
[0522] Step 4:
[0523] The terminal converts the extracted skeletal features into a standard format and sends them to the server. The data is encrypted before transmission to ensure security.
[0524] Step 5:
[0525] The server analyzes the user's voice and facial expressions simultaneously with the image data, and uses an emotion engine to recognize the user's emotional state. The recognized emotion data is then used for subsequent analysis.
[0526] Step 6:
[0527] The server compares the accumulated athlete data with the received skeletal features to determine the most suitable sport. This process utilizes machine learning algorithms to make the optimal selection based on skeletal attributes.
[0528] Step 7:
[0529] The server receives the results from the emotion engine and adjusts the recommended sports list according to the user's emotional state. For example, if the user is feeling stressed, sports with relaxing effects will be prioritized.
[0530] Step 8:
[0531] The server formats the adjusted sports recommendations and selection reasons for end users and sends them to their devices. The recommendations are customized to take into account the user's current mood.
[0532] Step 9:
[0533] The device decrypts the received data and displays it to the user in an easy-to-understand visual format. Based on this information, the user can make decisions about planning their child's sports activities.
[0534] (Example 2)
[0535] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0536] Currently, many systems suggest general exercises and activities to users, but these do not take into account the user's physical characteristics or emotional state. Therefore, they cannot suggest exercises that are optimal for individual needs, which can hinder user satisfaction and the ability to maintain effective activity. The challenge is to solve this problem and provide more personalized exercise recommendations.
[0537] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0538] In this invention, the server includes means for preprocessing image data, means for extracting skeletal features, and means for recognizing the user's emotional information. This enables personalized exercise recommendations based on the user's physical characteristics and emotional state.
[0539] "Image data" refers to data that represents visual information in a digital format, such as photographs and videos that are processed within a system.
[0540] "Preprocessing" refers to the initial stages of processing performed to make data easier to analyze, including tasks such as formatting standardization, noise reduction, and brightness adjustment.
[0541] "Skeletal features" refer to information about the positional relationships of joints and bones in the human body, extracted from image data.
[0542] "Adjustment" refers to the act of appropriately modifying data or results according to specific conditions or criteria, and in this context, it means optimizing exercise recommendations based on the user's emotional state.
[0543] The embodiment for carrying out the present invention is configured as a system that allows users to receive more personalized exercise suggestions.
[0544] Users run a dedicated application on their smartphone or tablet to acquire and upload image data of their child's entire body. This device has a function to preprocess the image data, performing image resizing, noise reduction, and brightness adjustment. This improves image quality and allows for accurate extraction of skeletal features.
[0545] The terminal extracts skeletal features from images using a computer vision algorithm. This algorithm, for example, uses an open-source computer vision library to identify joint points in a person. The extracted skeletal feature data is sent from the terminal to the server.
[0546] The server compares the received skeletal feature data with already stored human data to determine the optimal movement. A machine learning model runs on the server, rapidly optimizing based on the received data. Furthermore, the server integrates an emotion engine that analyzes the user's voice, facial expressions, and biometric data to extract emotional information. This analysis utilizes common speech recognition systems and facial expression analysis software.
[0547] The server adjusts the exercise recommendation list based on the user's emotional state. For example, it can broaden the exercise options if the user is relaxed, or prioritize specific exercises if they are stressed. This personalized exercise recommendation, along with the reasons for its selection, is ultimately sent to the user's device for visual confirmation.
[0548] For example, when the user is relaxed, the system is expected to recommend "extensive and challenging exercise." Conversely, if a state of tension is detected, it will recommend "easy and low-stress exercise." An example of a prompt to the generative AI model would be: "I would like to upload image data of my child and receive a sports recommendation. Please extract skeletal features and suggest the most suitable sport based on the user's emotional state. The user is currently relaxed."
[0549] This structure enables the present invention to recommend exercises that take into account the user's physical characteristics and emotional state, thereby supporting more effective and personalized health promotion.
[0550] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0551] Step 1:
[0552] The user uses a dedicated application to acquire image data of the child's entire body. The input is image data acquired from the camera, and the output is an image file saved on the device. This process prepares the basic data that the system will use for subsequent processing.
[0553] Step 2:
[0554] The device performs preprocessing on the acquired image data. The input is the image data acquired by the user, and the output is preprocessed, clean, and uniform image data. Specifically, the device performs image resizing, noise reduction, and brightness adjustment. This process allows for more effective and accurate extraction of skeletal features.
[0555] Step 3:
[0556] The device uses computer vision algorithms to extract skeletal features. The input is pre-processed image data, and the output is data on the joint positions of the identified skeleton. Specifically, it uses a library to detect points of the body within the image. The output obtained in this step becomes the data necessary for motion optimization.
[0557] Step 4:
[0558] The terminal sends extracted skeletal feature data to the server. The input is skeletal joint position data, and the output is data securely transferred to the server. A communication protocol is used to encode and compress the data, achieving efficient and secure data transfer.
[0559] Step 5:
[0560] The server compares the received skeletal feature data with a stored human database. The input is the transferred skeletal feature data, and the output is a movement suggestion based on the person's motor skills. A machine learning algorithm is used to analyze the data and determine which movement is appropriate.
[0561] Step 6:
[0562] The server uses an emotion engine to analyze the user's emotional information. Inputs include voice, facial expressions, and biometric data, while output is metadata indicating emotion. Emotion recognition software is used to identify emotions in real time, and the system adjusts accordingly.
[0563] Step 7:
[0564] The server adjusts exercise suggestions based on emotional information. Inputs are exercise suggestions and emotional metadata, while output is an exercise suggestion adapted to the user's emotions. An adjustment algorithm is executed to select an exercise that matches the user's current emotional state. This process enables the server to provide emotionally appropriate exercise guidance to the user.
[0565] Step 8:
[0566] The server sends the final exercise suggestions and the reasons for their selection to the terminal. The input is the adjusted list of exercise suggestions, and the output is the information displayed on the user's terminal. Data transfer technology is used to send the information to the terminal and display it in the most optimal format. As a result, the user receives specific exercise suggestions and can use them to implement them.
[0567] (Application Example 2)
[0568] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0569] In industrial environments, worker fatigue and emotional state can affect work efficiency and safety. Therefore, there is a need for methods to improve safety by monitoring workers' conditions in real time and providing work instructions tailored to their specific situations.
[0570] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0571] In this invention, the server includes means for acquiring image information, means for extracting skeletal features, and means for analyzing emotional states. This makes it possible to grasp the health status of workers in real time and provide appropriate work instructions.
[0572] "Image information" refers to information recorded in visual data format, including the images that are the subject of analysis.
[0573] "Skeletal features" refer to feature data extracted from image information that indicates the skeletal structure of a person or object.
[0574] "Athlete data" refers to a database of skeletal and related information of athletes accumulated over time, and is used as a standard for comparison.
[0575] A "user" is an individual or organization that obtains information or makes decisions through the system.
[0576] "Emotional state" refers to information that indicates an individual's current psychological and physiological state.
[0577] "Confidentiality" is the process of protecting information using specific technologies to prevent its misuse or leakage.
[0578] To realize this invention, a system is configured in which a server, terminal, and user work together. Specifically, the terminal is equipped with a high-resolution camera and microphone to acquire image information and audio data of the worker. The image information is preprocessed using an image processing library such as OpenCV to extract skeletal features. The extracted skeletal data is sent to the server and compared with previously accumulated data on athletes.
[0579] The server uses machine learning models built with PyTorch and TensorFlow to analyze the user's emotional state. Through data analysis, the server identifies the worker's fatigue level and stress level, and then derives optimal work instructions accordingly. If the server determines the user is relaxed, a wide range of tasks are suggested. Conversely, if tension is detected, calmer tasks are prioritized.
[0580] As a concrete example, factory workers are monitored by cameras and biosensors, and if fatigue is detected during work, the system prompts them to take appropriate breaks. It also recommends that workers switch from tasks requiring caution to safer tasks.
[0581] An example of a prompt for a generative AI model is: "Describe a program that uses video and audio data of a worker to analyze their skeletal structure and emotions, and then provides the worker with appropriate work instructions."
[0582] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0583] Step 1:
[0584] The terminal uses its built-in high-resolution camera and microphone to acquire image information and audio data of the worker. The input consists of camera images and audio data, which serve as the initial input data to the system. The output is sent for preprocessing in its original format.
[0585] Step 2:
[0586] The device performs preprocessing on the acquired image information using OpenCV, denoising and normalizing the image data. It also performs calculations to extract skeletal features from the image data. The input is raw image data, and the output is processed skeletal feature data.
[0587] Step 3:
[0588] The terminal sends skeletal feature data to the server. The server compares the received skeletal feature data with accumulated data on athletes. The input is skeletal feature data, and through comparison calculations, the optimal task list for the worker is generated. The output is the optimal task list.
[0589] Step 4:
[0590] The server uses PyTorch or TensorFlow to perform emotion analysis using voice and other sensor data. Inputs are voice data and, if necessary, sensor data. Through this analysis, the server determines and outputs the worker's emotional state.
[0591] Step 5:
[0592] The server integrates skeletal feature comparison results and sentiment analysis results to customize tasks to suit the worker. The input is the optimal task list and sentiment analysis results, and adjusted work instructions are generated based on this. The output is the adjusted work instruction list.
[0593] Step 6:
[0594] The adjusted work order list is sent to the terminal and displayed to the user. Based on this information, the user can perform or modify the work. The input is the adjusted work order list, and the output is the work based on the user's decisions.
[0595] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0596] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0597] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0598] [Fourth Embodiment]
[0599] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0600] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0601] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0602] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0603] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0604] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0605] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0606] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0607] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0608] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0609] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0610] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0611] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0612] To implement this invention, a user acquires image data of a child's entire body using a dedicated application. Upon receiving this image data, the terminal first performs preprocessing, such as noise reduction and size adjustment. Subsequently, it extracts skeletal features from the image using computer vision technology. Specifically, it identifies the positions of the face, shoulders, elbows, knees, and ankles, and calculates their relative distances and angles.
[0613] The extracted skeletal features are sent to the server in a standardized format. Upon receiving these skeletal features, the server compares them to a database of athletes and applies a machine learning algorithm to determine the most suitable sport. The server refers to the profiles of numerous athletes stored in the database and identifies sports with statistically similar skeletal characteristics. Based on this analysis, it generates a list of the most suitable sports and the reasons for their selection, and sends it to the user's terminal.
[0614] For example, suppose an analysis of image data of a specific child reveals that the child is tall and has long arms. The server determines that these characteristics are highly correlated with sports where they are typically advantageous, such as basketball, volleyball, and certain track and field events. It then generates a recommendation list that includes these sports and sends it to the terminal with an explanation such as "height and arm length are related."
[0615] In this way, the system of the present invention provides a means to enable users to make scientifically and objectively rational choices.
[0616] Based on the information received, users can assess their child's aptitude and create a concrete plan for sports activities. This helps unlock the child's potential and provides them with opportunities to enjoy sports in an optimal environment.
[0617] The following describes the processing flow.
[0618] Step 1:
[0619] The user uploads image data of the child's entire body to a dedicated application. The user ensures that the child is sitting upright when the photo is taken.
[0620] Step 2:
[0621] The device preprocesses the image data it receives. Specifically, it adjusts the image resolution and removes noise as needed to prepare the image for analysis.
[0622] Step 3:
[0623] The device uses a computer vision algorithm to extract skeletal features from image data. The algorithm identifies the location of each point in the skeleton, such as the shoulders, elbows, knees, and ankles, and calculates the distances and angles between these locations.
[0624] Step 4:
[0625] The terminal converts the extracted skeletal features into a standardized format, encrypts them for security, and then sends them to the server.
[0626] Step 5:
[0627] The server analyzes the skeletal features it receives and compares them to accumulated athlete data. Here, a machine learning algorithm is used to identify the sport with the highest correlation.
[0628] Step 6:
[0629] Based on the analysis results, the server generates a list of the most suitable sports for the child, along with the reasons for their selection. This includes explaining why certain skeletal characteristics are advantageous for specific sports.
[0630] Step 7:
[0631] The server encrypts the generated sports recommendation list and sends it to the terminal. Privacy protection is taken into consideration during this process.
[0632] Step 8:
[0633] The device decrypts the received data and displays it in an easy-to-understand format for the user. Based on this, the user can then plan their child's sports activities.
[0634] (Example 1)
[0635] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0636] Conventional sports aptitude assessment systems suffered from insufficient preprocessing, data transmission, and analysis methods in the process from image data acquisition to aptitude assessment, resulting in challenges in terms of prediction accuracy and security. Furthermore, the feedback provided to users before suggesting appropriate exercise activities was inadequate, sometimes failing to deliver reliable judgments.
[0637] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0638] In this invention, the server includes means for preprocessing image information, means for transmitting joint features in a standard format, and means for analyzing the data using a clustering algorithm. This enables effective and reliable assessment of suitability for exercise activities by statistically comparing it with data from athletes.
[0639] "Image information" refers to visual data acquired in digital format that includes the visual characteristics of a specific object.
[0640] "Preprocessing" refers to the process of applying noise reduction, size adjustment, and other adjustments to acquired image information to prepare the data for a useful format.
[0641] "Joint features" refer to data extracted from image information that shows the location of joints in the human body and their spatial relationships.
[0642] A "standard format" is a structured data format that is consistently used in the transmission, reception, and processing of data.
[0643] A "clustering algorithm" refers to a computational method used to statistically analyze data and group elements that have similar characteristics.
[0644] "Athlete data" refers to a collection of information regarding the physical characteristics and performance of athletes engaged in diverse athletic activities.
[0645] "Assessing suitability for physical activity" is the act of identifying the most suitable exercise or sport for an individual based on analyzed data.
[0646] To implement this invention, the user first uses a dedicated application to acquire full-body image information of a child using a camera-equipped device. The terminal then performs preprocessing on the received image information, such as noise reduction and size adjustment, using an image processing library. Specifically, the terminal utilizes image processing software such as OpenCV to optimize the image quality.
[0647] Next, the device utilizes computer vision technology to extract joint features from the image. This process uses OpenPose, a pre-trained model, to calculate the coordinates of the face and joint positions. This information is structured in a standard format such as JSON and treated as data.
[0648] The terminal sends standardized joint feature data to the server using the HTTPS protocol. After receiving this data, the server applies a clustering algorithm to compare it with the accumulated data on athletes. The server uses machine learning libraries such as Scikit-learn and TensorFlow, employing techniques such as K-means clustering, to identify sports with features similar to the input data.
[0649] Based on this analysis, the server generates a list of suitable sports and the reasons for their selection, and sends the information back to the user's device. The user can then review these results on their device and plan exercise activities that are tailored to their child's characteristics.
[0650] For example, if an analysis of a child's image data reveals features such as being tall and having long arms, the server will determine that these features are advantageous in sports like basketball or volleyball. This determination is then returned to the user with a reason for the selection, such as "because height and arm length are correlated."
[0651] An example of a prompt statement can be expressed as follows:
[0652] "Please provide information on a system that suggests optimal exercise activities based on joint features identified from images of children. Specifically, I'd like to know which features correlate with which sports, and how the server analyzes and returns the results."
[0653] Thus, the system of the present invention provides a means to support sports selection in a scientific and objective manner.
[0654] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0655] Step 1:
[0656] The user uses a dedicated application to capture full-body images of their child using the camera on their smartphone or tablet. The images are saved in JPEG or PNG format. The captured image information is then transmitted to the device through the application's interface.
[0657] Step 2:
[0658] The terminal performs preprocessing on the received image information. It receives an image in JPEG or PNG format as input, uses an image processing library to remove noise (using a Gaussian filter) and adjust the size (resizing to 224x224 pixels), and generates a clean, appropriately sized image as output.
[0659] Step 3:
[0660] The device extracts joint features using pre-processed images. Using pre-processed images as input, it detects the positions of faces and joints using computer vision techniques such as OpenPose. As output, it generates a dataset containing coordinate information for each joint.
[0661] Step 4:
[0662] The terminal converts the extracted joint features into a standard format (such as JSON). It uses joint coordinate information as input and converts this data into a structured format such as JSON. As output, it generates a data package that can be sent to the server.
[0663] Step 5:
[0664] The terminal sends a data package of joint features to the server. Using the HTTPS protocol, the data is securely transferred to the server, and once the transmission is complete, the terminal enters a state of waiting for a response from the server.
[0665] Step 6:
[0666] The server performs analysis based on the received joint feature data. Using the received coordinate data as input, it analyzes the data using algorithms such as K-means clustering, utilizing machine learning libraries (Scikit-learn, TensorFlow, etc.). As output, it generates a list of exercise activities suitable for the subject.
[0667] Step 7:
[0668] The server generates analysis results (a list of suitable exercise activities and the reasons for them), encrypts them, and sends them to the terminal. The output provides data in a format viewable on the user's terminal.
[0669] Step 8:
[0670] Users receive feedback from the server on their device and review a list of suitable sports. Based on this information, they can plan their child's sports activities.
[0671] (Application Example 1)
[0672] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0673] In delivery services, a challenge is that delivery personnel are unable to select the most efficient delivery areas and routes based on their individual physical characteristics. As a result, they may not be able to make the most of their individual abilities, potentially leading to decreased work efficiency. This invention aims to eliminate this inefficiency and provide a system that proposes an appropriate delivery process for delivery personnel.
[0674] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0675] In this invention, the server includes means for acquiring image data, means for extracting physical characteristics from the image data, means for comparing the physical characteristics with pre-stored individual data to determine the optimal behavioral pattern, means for presenting the determination result to the user, and means for recommending an efficient region or route based on the determination result. This makes it possible to suggest the optimal delivery region and route for each delivery person.
[0676] "Image data" refers to a record of the visual information of an object in numerical format.
[0677] "Physical characteristics" refer to data that describes an individual's skeletal structure and physical features.
[0678] "Individual data" refers to a collection of information that accumulates the physical characteristics of multiple individuals.
[0679] A "behavioral pattern" is a pattern of optimal behavior or activity that is predicted based on an individual's characteristics.
[0680] "Users" refer to entities or individuals that utilize this system.
[0681] "Region or route" refers to the specific locations or travel routes where the delivery person performs their duties.
[0682] "Recommendation" means presenting the option that best suits specific conditions.
[0683] In implementing this invention, the user first uses a terminal such as a smartphone to acquire whole-body image data of the target individual. The terminal utilizes advanced image processing techniques to extract physical characteristics from the image data. For this processing, OpenCV, an open-source computer vision library, is used for noise reduction and size adjustment.
[0684] The physical characteristics data extracted by the device is sent to the server. The server compares the received physical characteristics with existing data in the individual database and uses a generative AI model to determine the optimal behavioral pattern. During this process, the server performs data analysis using machine learning libraries such as TensorFlow and PyTorch.
[0685] The results are presented to the user. These results include the most suitable region or route for the individual, and suggest efficient actions and activities.
[0686] For example, if a delivery person has physical characteristics such as being tall and having long arms, the server will determine that they are suitable for long-distance deliveries and suggest a delivery route that covers a wide area. This suggestion is then notified to the user's smartphone.
[0687] An example of a prompt message to be input into the generating AI model is: "Generate an efficient delivery route based on the physical characteristics of the delivery person. For example, for a tall delivery person with long arms, suggest a route that covers a wide area." In this way, the present invention provides a system that enables efficient work suited to the characteristics of the delivery person.
[0688] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0689] Step 1:
[0690] The user uses their smartphone camera to take a full-body image of the delivery person. The input is the captured full-body image data, and the output is an image file stored on the device. The device starts processing this image data as its initial input.
[0691] Step 2:
[0692] The terminal uses the OpenCV library to preprocess the acquired image data. The input is whole-body image data, and filtering and resizing are performed to remove noise and adjust the size of the image. The output is preprocessed, clear image data. Specifically, the image sharpness is improved and the resolution is adjusted to facilitate analysis.
[0693] Step 3:
[0694] The device uses TensorFlow to extract body characteristics based on pre-processed image data. The input is pre-processed image data, and the output is body characteristic data such as the position and length of the skeleton. This data extraction uses a pre-trained model to identify the positions of the face, shoulders, elbows, knees, and ankles, and to calculate their relative distances and angles.
[0695] Step 4:
[0696] The terminal sends extracted physical characteristic data to the server. The server receives this data and compares it with existing data in the individual database. The input is physical characteristic data, and the output is identification information for similar database entries. Using a database search algorithm, the server applies a statistical model to efficiently determine similarity.
[0697] Step 5:
[0698] The server runs a machine learning model and compares it with an individual database to determine the optimal behavioral pattern. The input is the physical characteristics data received by the server and the information in the database, and the output is a suggestion of the optimal region or route specific to the individual. This model uses a deep learning algorithm implemented in TensorFlow or PyTorch.
[0699] Step 6:
[0700] The server sends the determined suggestion to the user. The input is the region or route information determined by the server, and the output is a notification message displayed on the user's device. The notification includes a brief explanation of the delivery route details and efficiency. Specifically, the notification message appears as a pop-up on the user's smartphone.
[0701] Step 7:
[0702] Users review notifications and perform delivery tasks according to the suggested area or route. The input is the information presented as a notification, and the output is the streamlined actual delivery task. By adhering to these instructions, users can expect improved operational efficiency and minimized effort.
[0703] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0704] This invention begins with a user using a dedicated application to acquire image data of a child's entire body and uploading this data to a terminal. The image data is preprocessed on the terminal and processed by a computer vision algorithm to extract skeletal features. After detecting the skeletal features, the terminal sends the feature data to a server.
[0705] The server uses the received skeletal features to compare them with pre-stored athlete data and determine the most suitable sport. Furthermore, this system incorporates an emotion engine that recognizes the user's emotions when using the system. The emotion engine analyzes and recognizes the user's emotions from inputs such as voice, facial expressions, and biosensors.
[0706] Based on information obtained from the emotion engine, the server adjusts the sports recommendation results considering the user's current emotional state. In this process, the server uses the emotion recognition results to customize the recommendations according to the user's preferences and reactions, generating an optimized sports recommendation list for each user.
[0707] For example, if the user's voice and facial expressions indicate a relaxed state, the server could broaden the range of sports suggested, offering a wider variety of options. Conversely, if tension is detected, the server could narrow down the suggestions or prioritize sports considered to be less mentally taxing.
[0708] Finally, the server-generated list of adaptive recommendations based on emotions, along with the reasons for their selection, is sent to the device. The device displays this information clearly to the user and provides feedback. In this way, a more personalized selection of sports, taking emotional considerations into account, becomes possible. Users can use these results to plan their child's sports activities and receive information to help their child enjoy sports more.
[0709] The following describes the processing flow.
[0710] Step 1:
[0711] The user uploads image data of the child's entire body to a dedicated application. The user pays attention to whether the child is relaxed when taking the picture, and also records their facial expressions and voice.
[0712] Step 2:
[0713] The terminal processes the received image data, performing pre-processing such as noise reduction and image size adjustment. This pre-processing prepares the image for analysis.
[0714] Step 3:
[0715] The device applies computer vision algorithms to extract skeletal features from images. It identifies the locations of the face, shoulders, elbows, knees, ankles, etc., and calculates the associated skeletal dimensions.
[0716] Step 4:
[0717] The terminal converts the extracted skeletal features into a standard format and sends them to the server. The data is encrypted before transmission to ensure security.
[0718] Step 5:
[0719] The server analyzes the user's voice and facial expressions simultaneously with the image data, and uses an emotion engine to recognize the user's emotional state. The recognized emotion data is then used for subsequent analysis.
[0720] Step 6:
[0721] The server compares the accumulated athlete data with the received skeletal features to determine the most suitable sport. This process utilizes machine learning algorithms to make the optimal selection based on skeletal attributes.
[0722] Step 7:
[0723] The server receives the results from the emotion engine and adjusts the recommended sports list according to the user's emotional state. For example, if the user is feeling stressed, sports with relaxing effects will be prioritized.
[0724] Step 8:
[0725] The server formats the adjusted sports recommendations and selection reasons for end users and sends them to their devices. The recommendations are customized to take into account the user's current mood.
[0726] Step 9:
[0727] The device decrypts the received data and displays it to the user in an easy-to-understand visual format. Based on this information, the user can make decisions about planning their child's sports activities.
[0728] (Example 2)
[0729] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0730] Currently, many systems suggest general exercises and activities to users, but these do not take into account the user's physical characteristics or emotional state. Therefore, they cannot suggest exercises that are optimal for individual needs, which can hinder user satisfaction and the ability to maintain effective activity. The challenge is to solve this problem and provide more personalized exercise recommendations.
[0731] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0732] In this invention, the server includes means for preprocessing image data, means for extracting skeletal features, and means for recognizing the user's emotional information. This enables personalized exercise recommendations based on the user's physical characteristics and emotional state.
[0733] "Image data" refers to data that represents visual information in a digital format, such as photographs and videos that are processed within a system.
[0734] "Preprocessing" refers to the initial stages of processing performed to make data easier to analyze, including tasks such as formatting standardization, noise reduction, and brightness adjustment.
[0735] "Skeletal features" refer to information about the positional relationships of joints and bones in the human body, extracted from image data.
[0736] "Adjustment" refers to the act of appropriately modifying data or results according to specific conditions or criteria, and in this context, it means optimizing exercise recommendations based on the user's emotional state.
[0737] The embodiment for carrying out the present invention is configured as a system that allows users to receive more personalized exercise suggestions.
[0738] Users run a dedicated application on their smartphone or tablet to acquire and upload image data of their child's entire body. This device has a function to preprocess the image data, performing image resizing, noise reduction, and brightness adjustment. This improves image quality and allows for accurate extraction of skeletal features.
[0739] The terminal extracts skeletal features from images using a computer vision algorithm. This algorithm, for example, uses an open-source computer vision library to identify joint points in a person. The extracted skeletal feature data is sent from the terminal to the server.
[0740] The server compares the received skeletal feature data with already stored human data to determine the optimal movement. A machine learning model runs on the server, rapidly optimizing based on the received data. Furthermore, the server integrates an emotion engine that analyzes the user's voice, facial expressions, and biometric data to extract emotional information. This analysis utilizes common speech recognition systems and facial expression analysis software.
[0741] The server adjusts the exercise recommendation list based on the user's emotional state. For example, it can broaden the exercise options if the user is relaxed, or prioritize specific exercises if they are stressed. This personalized exercise recommendation, along with the reasons for its selection, is ultimately sent to the user's device for visual confirmation.
[0742] For example, when the user is relaxed, the system is expected to recommend "extensive and challenging exercise." Conversely, if a state of tension is detected, it will recommend "easy and low-stress exercise." An example of a prompt to the generative AI model would be: "I would like to upload image data of my child and receive a sports recommendation. Please extract skeletal features and suggest the most suitable sport based on the user's emotional state. The user is currently relaxed."
[0743] This structure enables the present invention to recommend exercises that take into account the user's physical characteristics and emotional state, thereby supporting more effective and personalized health promotion.
[0744] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0745] Step 1:
[0746] The user uses a dedicated application to acquire image data of the child's entire body. The input is image data acquired from the camera, and the output is an image file saved on the device. This process prepares the basic data that the system will use for subsequent processing.
[0747] Step 2:
[0748] The device performs preprocessing on the acquired image data. The input is the image data acquired by the user, and the output is preprocessed, clean, and uniform image data. Specifically, the device performs image resizing, noise reduction, and brightness adjustment. This process allows for more effective and accurate extraction of skeletal features.
[0749] Step 3:
[0750] The device uses computer vision algorithms to extract skeletal features. The input is pre-processed image data, and the output is data on the joint positions of the identified skeleton. Specifically, it uses a library to detect points of the body within the image. The output obtained in this step becomes the data necessary for motion optimization.
[0751] Step 4:
[0752] The terminal sends extracted skeletal feature data to the server. The input is skeletal joint position data, and the output is data securely transferred to the server. A communication protocol is used to encode and compress the data, achieving efficient and secure data transfer.
[0753] Step 5:
[0754] The server compares the received skeletal feature data with a stored human database. The input is the transferred skeletal feature data, and the output is a movement suggestion based on the person's motor skills. A machine learning algorithm is used to analyze the data and determine which movement is appropriate.
[0755] Step 6:
[0756] The server uses an emotion engine to analyze the user's emotional information. Inputs include voice, facial expressions, and biometric data, while output is metadata indicating emotion. Emotion recognition software is used to identify emotions in real time, and the system adjusts accordingly.
[0757] Step 7:
[0758] The server adjusts exercise suggestions based on emotional information. Inputs are exercise suggestions and emotional metadata, while output is an exercise suggestion adapted to the user's emotions. An adjustment algorithm is executed to select an exercise that matches the user's current emotional state. This process enables the server to provide emotionally appropriate exercise guidance to the user.
[0759] Step 8:
[0760] The server sends the final exercise suggestions and the reasons for their selection to the terminal. The input is the adjusted list of exercise suggestions, and the output is the information displayed on the user's terminal. Data transfer technology is used to send the information to the terminal and display it in the most optimal format. As a result, the user receives specific exercise suggestions and can use them to implement them.
[0761] (Application Example 2)
[0762] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0763] In industrial environments, worker fatigue and emotional state can affect work efficiency and safety. Therefore, there is a need for methods to improve safety by monitoring workers' conditions in real time and providing work instructions tailored to their specific situations.
[0764] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0765] In this invention, the server includes means for acquiring image information, means for extracting skeletal features, and means for analyzing emotional states. This makes it possible to grasp the health status of workers in real time and provide appropriate work instructions.
[0766] "Image information" refers to information recorded in visual data format, including the images that are the subject of analysis.
[0767] "Skeletal features" refer to feature data extracted from image information that indicates the skeletal structure of a person or object.
[0768] "Athlete data" refers to a database of skeletal and related information of athletes accumulated over time, and is used as a standard for comparison.
[0769] A "user" is an individual or organization that obtains information or makes decisions through the system.
[0770] "Emotional state" refers to information that indicates an individual's current psychological and physiological state.
[0771] "Confidentiality" is the process of protecting information using specific technologies to prevent its misuse or leakage.
[0772] To realize this invention, a system is configured in which a server, terminal, and user work together. Specifically, the terminal is equipped with a high-resolution camera and microphone to acquire image information and audio data of the worker. The image information is preprocessed using an image processing library such as OpenCV to extract skeletal features. The extracted skeletal data is sent to the server and compared with previously accumulated data on athletes.
[0773] The server uses machine learning models built with PyTorch and TensorFlow to analyze the user's emotional state. Through data analysis, the server identifies the worker's fatigue level and stress level, and then derives optimal work instructions accordingly. If the server determines the user is relaxed, a wide range of tasks are suggested. Conversely, if tension is detected, calmer tasks are prioritized.
[0774] As a concrete example, factory workers are monitored by cameras and biosensors, and if fatigue is detected during work, the system prompts them to take appropriate breaks. It also recommends that workers switch from tasks requiring caution to safer tasks.
[0775] An example of a prompt for a generative AI model is: "Describe a program that uses video and audio data of a worker to analyze their skeletal structure and emotions, and then provides the worker with appropriate work instructions."
[0776] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0777] Step 1:
[0778] The terminal uses its built-in high-resolution camera and microphone to acquire image information and audio data of the worker. The input consists of camera images and audio data, which serve as the initial input data to the system. The output is sent for preprocessing in its original format.
[0779] Step 2:
[0780] The device performs preprocessing on the acquired image information using OpenCV, denoising and normalizing the image data. It also performs calculations to extract skeletal features from the image data. The input is raw image data, and the output is processed skeletal feature data.
[0781] Step 3:
[0782] The terminal sends skeletal feature data to the server. The server compares the received skeletal feature data with accumulated data on athletes. The input is skeletal feature data, and through comparison calculations, the optimal task list for the worker is generated. The output is the optimal task list.
[0783] Step 4:
[0784] The server uses PyTorch or TensorFlow to perform emotion analysis using voice and other sensor data. Inputs are voice data and, if necessary, sensor data. Through this analysis, the server determines and outputs the worker's emotional state.
[0785] Step 5:
[0786] The server integrates skeletal feature comparison results and sentiment analysis results to customize tasks to suit the worker. The input is the optimal task list and sentiment analysis results, and adjusted work instructions are generated based on this. The output is the adjusted work instruction list.
[0787] Step 6:
[0788] The adjusted work order list is sent to the terminal and displayed to the user. Based on this information, the user can perform or modify the work. The input is the adjusted work order list, and the output is the work based on the user's decisions.
[0789] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0790] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0791] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0792] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0793] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0794] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0795] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0796] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0797] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0798] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0799] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0800] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0801] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0802] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0803] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0804] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0805] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0806] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0807] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0808] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0809] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0810] The following is further disclosed regarding the embodiments described above.
[0811] (Claim 1)
[0812] Means for acquiring image data,
[0813] A means for extracting skeletal features from the image data,
[0814] A means for determining a suitable sport by comparing the skeletal characteristics with pre-stored athlete data,
[0815] A means of presenting the result of the judgment to the user,
[0816] A system that includes this.
[0817] (Claim 2)
[0818] The system according to claim 1, further comprising means for preprocessing the image data.
[0819] (Claim 3)
[0820] The system according to claim 1, further comprising means for encrypting and transmitting the determination result.
[0821] "Example 1"
[0822] (Claim 1)
[0823] Means for acquiring image information,
[0824] Means for preprocessing the image information,
[0825] A means for extracting joint features from the image information,
[0826] Means for transmitting the joint features in a standard format,
[0827] A means for determining appropriate exercise activities by comparing the joint characteristics with pre-stored data on athletes,
[0828] A means of presenting the result of the judgment to the user,
[0829] A system that includes this.
[0830] (Claim 2)
[0831] The system according to claim 1, further comprising means for analyzing the joint features using a clustering algorithm.
[0832] (Claim 3)
[0833] The system according to claim 1, further comprising means for securely transmitting the determination result.
[0834] "Application Example 1"
[0835] (Claim 1)
[0836] Means for acquiring image data,
[0837] A means for extracting physical characteristics from the image data,
[0838] A means for determining the optimal behavioral pattern by comparing the physical characteristics with pre-accumulated individual data,
[0839] A means of presenting the result of the judgment to the user,
[0840] A means of recommending an efficient region or route based on the results of the determination,
[0841] A system that includes this.
[0842] (Claim 2)
[0843] The system according to claim 1, further comprising means for preprocessing the image data.
[0844] (Claim 3)
[0845] The system according to claim 1, further comprising means for encrypting and transmitting the determination result.
[0846] "Example 2 of combining an emotion engine"
[0847] (Claim 1)
[0848] Means for acquiring image data,
[0849] Means for preprocessing the image data,
[0850] A means for extracting skeletal features from the image data,
[0851] A means for determining appropriate exercise by comparing the skeletal characteristics with accumulated data on a person,
[0852] A means for adjusting the judgment result based on the user's emotional state,
[0853] A means for presenting the adjusted result to the user,
[0854] A system that includes this.
[0855] (Claim 2)
[0856] The system according to claim 1, further comprising means for recognizing the user's emotional information using voice, facial expression, and biometric inputs in the determination.
[0857] (Claim 3)
[0858] The system according to claim 1, further comprising means for transmitting the judgment result and the adjusted recommendation result.
[0859] "Application example 2 when combining with an emotional engine"
[0860] (Claim 1)
[0861] Means for acquiring image information,
[0862] A means for extracting skeletal features from the image information,
[0863] A means for determining suitable exercises by comparing the skeletal characteristics with accumulated data on athletes,
[0864] A means of presenting the result of the judgment to the user,
[0865] A means of analyzing the emotional state of users,
[0866] Means for adjusting the judgment result based on the emotional state,
[0867] A system that includes this.
[0868] (Claim 2)
[0869] The system according to claim 1, further comprising means for preprocessing the image information.
[0870] (Claim 3)
[0871] The system according to claim 1, further comprising means for transmitting the determination result in a confidential manner. [Explanation of symbols]
[0872] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for acquiring image data, A means for extracting skeletal features from the image data, A means for determining a suitable sport by comparing the skeletal characteristics with pre-stored athlete data, A means of presenting the result of the judgment to the user, A system that includes this.
2. The system according to claim 1, further comprising means for preprocessing the image data.
3. The system according to claim 1, further comprising means for encrypting and transmitting the determination result.