System
The system addresses bodybuilding competition challenges by using AI for real-time skin color correction and muscle evaluation, ensuring fair assessments and enhancing viewer engagement.
Patent Information
- Application Number
- JP2024118092
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
Athletes in bodybuilding competitions face challenges with skin darkening to enhance muscle prominence, leading to time and cost inefficiencies, skin fairness issues, and contamination risks, which hinder performance and fair evaluation.
A system utilizing artificial intelligence for real-time skin color correction, muscle evaluation, and distribution of corrected video and scores to ensure fair muscle assessment.
Enables fair and transparent muscle evaluation in real-time, eliminating the need for skin darkening and reducing venue contamination, while providing convenient and entertaining viewing experiences.
Smart Images

Figure 2026017310000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In bodybuilding competitions, athletes are required to unnecessarily darken their skin to make their muscles more prominent, which poses a major challenge in terms of the time and cost involved. Furthermore, skin color can lead to a lack of fairness, making it difficult for athletes to concentrate on their performance. Another problem is the contamination of the competition venue due to sunburn and makeup. A new method is needed to resolve these issues and enable fair muscle evaluation. [Means for solving the problem]
[0005] The present invention provides a system including a means for inputting video, a means for correcting skin color in the input video, a means for evaluating muscles based on the corrected video, and a means for delivering the muscle evaluation and corrected video in real time. Specifically, the skin color correction means uses an artificial intelligence model to unify skin tones, and the muscle evaluation means analyzes muscle shading and contours from the corrected video, enabling fair muscle evaluation. This system eliminates the need for athletes to darken their skin, providing an environment where they can concentrate on their performance. It also solves the problem of uncleanness at competition venues and provides viewers with fair evaluations in real time.
[0006] "Means for inputting video" refers to hardware and software for receiving video data from a camera or other video capture device.
[0007] The "skin color correction means" refers to an artificial intelligence model and algorithm for uniformly correcting skin color in the input video to unify skin tones.
[0008] A "muscle assessment means" is software and algorithms for analyzing muscle shadows and contours from the corrected image and generating a muscle assessment score.
[0009] "Means for real-time distribution" means hardware and software for distributing the corrected footage and muscle evaluation scores to viewers and judges in real time via the Internet or other communication means.
[0010] An "artificial intelligence model" is an algorithm and its implementation that uses technologies such as machine learning and deep learning to perform skin tone correction and muscle evaluation.
[0011] "Color correction" is an image processing technique for standardizing skin tones in input video.
[0012] Analyzing "muscle shading and contours" is a technology that detects the shape and shading of muscles in an image and calculates a muscle evaluation score based on that.
[0013] The "evaluation score" is a numerical value obtained from the analysis based on the shadows and contours of the muscles. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The present invention is a system for correcting the skin color of athletes in bodybuilding competitions and achieving fair muscle evaluation. The following describes the specific processing of each system element and program. The present invention is realized by the server, terminals, and users.
[0036] Server-side processing
[0037] 1. Acquiring footage
[0038] The server acquires video data from cameras and other video acquisition devices. The video is received in real-time streaming.
[0039] 2. Skin Tone Correction
[0040] The server analyzes the received video data and uses an AI model to correct the players' skin tones. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and processes them to unify their color tones.
[0041] 3. Muscle Assessment
[0042] The corrected image is further analyzed to detect muscle shadows and contours. The server then calculates a muscle evaluation score based on this. The muscle evaluation algorithm uses edge detection and depth analysis to evaluate the degree to which muscles stand out in detail.
[0043] 4. Real-time streaming
[0044] The server distributes the corrected video and muscle evaluation scores to viewers and judges in real time via the Internet or dedicated communication protocols (e.g., RTSP, WebSocket).
[0045] Terminal side processing
[0046] 1. Receiving video and evaluation data
[0047] The terminal receives the corrected video and muscle evaluation data transmitted from the server.
[0048] 2. Displaying images
[0049] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[0050] 3. Displaying evaluation data
[0051] The device displays the muscle evaluation score in real time, which is rendered as a number in a specific area using DOM manipulation.
[0052] User Actions
[0053] 1. Video selection
[0054] Users can select the video of the player they want to evaluate using the device, for example, a remote control or touch interface.
[0055] 2. Watching the video
[0056] Users can view the corrected footage and check the muscle evaluation score in real time. They can also switch to other athletes while viewing.
[0057] 3. Providing Feedback
[0058] Users can provide feedback about their viewing experience to the system, for example by entering and submitting comments and ratings using a rating form within the application.
[0059] Specific examples
[0060] Example of server operation: Video is acquired in real time, skin color is corrected using an AI model, and the corrected video is distributed along with muscle analysis results. The server repeats this process to accumulate evaluation data for each athlete.
[0061] Example of device operation: The corrected video and muscle evaluation score received from the server are displayed on the screen. When the user presses a button to switch to the video of another athlete, a different corrected video and evaluation score are displayed.
[0062] Example of user operation: Users use the application on their smartphone or tablet to select their favorite players to watch, and enjoy watching the game while referring to real-time evaluation data. After watching the game, they also submit feedback using the evaluation form in the app.
[0063] In this way, the present invention is a system that realizes a fair and transparent bodybuilding competition through video correction and muscle evaluation, thereby providing convenience and entertainment not only to athletes but also to viewers.
[0064] The processing flow will be explained below.
[0065] Server-side processing flow
[0066] Step 1:
[0067] The server acquires live video from cameras and video acquisition devices. The video signal is received in streaming format and is captured as video data in real time.
[0068] Step 2:
[0069] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[0070] Step 3:
[0071] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[0072] Step 4:
[0073] The corrected video data is then fed into another AI model to analyze muscle shading and contours. The server uses edge detection and shading analysis algorithms to assess muscle prominence.
[0074] Step 5:
[0075] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[0076] Step 6:
[0077] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[0078] Processing flow on the terminal side
[0079] Step 1:
[0080] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[0081] Step 2:
[0082] The device analyzes the received data and extracts the video data. <video>It is displayed on the screen using tags and canvases.
[0083] Step 3:
[0084] The device also receives muscle evaluation scores and displays them in a specific area on the screen. The evaluation scores are updated in real time through DOM manipulation.
[0085] User Operation Flow
[0086] Step 1:
[0087] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[0088] Step 2:
[0089] Users can view the retouched footage of their chosen athlete while viewing their muscle assessment score, which is updated in real time and displayed simultaneously.
[0090] Step 3:
[0091] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[0092] In this way, the server, terminals, and users each play their respective roles, and a system is operated that realizes a fair and transparent bodybuilding competition in real time.
[0093] Example 1
[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0095] In existing bodybuilding competitions, the skin color of athletes affects the evaluation, making it difficult to provide a fair muscle evaluation. Additionally, there is a lack of systems that allow judges and viewers to provide fair evaluations in real time. Therefore, an effective method to correct athletes' skin color and provide a fair muscle evaluation is needed.
[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0097] In this invention, the server includes means for acquiring video from a video acquisition device, means for correcting skin color in the acquired video, means for calculating a muscle evaluation score based on the corrected video, means for distributing the corrected video and muscle evaluation score in real time, means for displaying the received corrected video, and means for displaying the muscle evaluation score in real time. This allows for fair evaluation by correcting the skin color of the athlete, and makes it possible to provide fair muscle evaluations in real time to judges and viewers.
[0098] "Video capture device" refers to a camera or other video capture mechanism that captures video data in real time.
[0099] The "skin color correction means" is a technology that performs processing to ensure consistent skin color tones of players in the captured video.
[0100] The "means for calculating muscle evaluation scores" is a technology that analyzes the corrected video to evaluate the shading and contours of the player's muscles and convert them into a numerical score.
[0101] "Real-time distribution means" means technology that distributes the corrected footage and muscle evaluation scores to viewers and judges in real time without delay via the Internet or other communications protocols.
[0102] The "means for displaying corrected images" is a technique for displaying corrected images received from the server on the screen of the terminal.
[0103] "Means for displaying muscle evaluation scores in real time" refers to a technology that instantly displays muscle evaluation scores received from a server on the screen of a terminal.
[0104] A "generative AI model" is an artificial intelligence model that learns from large amounts of data and is designed to perform specific tasks.
[0105] The present invention is a system for correcting the skin color of athletes in bodybuilding competitions and achieving fair muscle evaluation. The present invention is realized by the server, terminals, and users.
[0106] Server-side processing
[0107] The server first acquires video data in real time from a video capture device, such as a high-resolution camera. The video is then sent to the server using the RTSP protocol, where it is received and processed by software such as FFmpeg.
[0108] The server then uses generative AI models such as TensorFlow to correct the skin tone of the captured video. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and uses techniques such as histogram matching to unify the skin tone.
[0109] The corrected video is analyzed for muscle shading and contours using OpenCV. The server performs this analysis using the Canny edge detection algorithm and depth analysis technology, and then calculates a muscle evaluation score using a proprietary algorithm based on the obtained data.
[0110] The server also uses WebSocket to deliver the corrected video and muscle evaluation scores to viewers and judges in real time. The muscle evaluation scores are overlaid on each frame, and viewers can view them in a browser or a dedicated app.
[0111] Terminal side processing
[0112] The device receives the corrected video and muscle evaluation data sent from the server using WebSocket as the reception protocol.
[0113] The device analyzes the received data and <video>The corrected image is displayed in real time using tags and Canvas technology, and the muscle evaluation score is displayed on the screen using JavaScript and DOM manipulation, allowing users to instantly check the corrected image and real-time muscle evaluation score.
[0114] User operations
[0115] The user selects the video of the desired player through the device. Using a remote control or touch interface, the desired player can be selected from a list of players on the screen. The video of the selected player is requested from the server, and the corrected video is sent to the device.
[0116] Users can check their muscle evaluation score in real time while watching the corrected video. After watching, they can provide feedback about their viewing experience using the in-app rating form. This feedback data is sent to the server and used for future analysis and improvements.
[0117] Specific examples
[0118] Example of server operation: The server acquires images in real time from a high-resolution camera and performs skin color correction using TensorFlow. Next, it analyzes muscle shading and contours using OpenCV, and delivers the corrected images and evaluation data in real time via WebSocket.
[0119] Example prompt sentence:
[0120] Capture video in real time, correct skin color using TensorFlow, analyze muscles using Canny edge detection, and deliver the results via WebSocket.
[0121] Example of device operation: The device receives the corrected video and muscle evaluation score from the server and sends them to the HTML5 <video>The tag and Canvas are used to display the player on the screen. When the user presses a button on the remote control to switch players, the new video and evaluation score are displayed.
[0122] Example prompt sentence:
[0123] Video received from the server is processed as HTML5 <video>Play with tags and use Canvas to display muscle evaluation scores. Switch players when the user selects them with the remote.
[0124] Example of user operation: The user uses the smartphone app to select their favorite player by touch operation and watch. After watching, they submit feedback using the evaluation form in the app.
[0125] Example prompt sentence:
[0126] Select a player by touching the smartphone app, and view the adjusted footage and muscle evaluation score. After viewing, submit your feedback in the evaluation form.
[0127] Thus, the present invention is a system that realizes fair and transparent bodybuilding competitions through video correction and muscle evaluation, and provides convenience and entertainment for both athletes and viewers.
[0128] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0129] Step 1:
[0130] The server acquires the video
[0131] Specifically, the server captures real-time video of the players using a high-resolution camera. This video data is sent to the server via the RTSP protocol. The input data is the raw video stream, and the output is the video data available within the server. The server receives and stores the video data using software such as FFmpeg.
[0132] Step 2:
[0133] The server corrects skin tones
[0134] After acquiring the video data, the server uses a generative AI model such as TensorFlow to perform skin tone correction. Specifically, the AI model receives the video data as input, recognizes skin-tone areas, and uses histogram matching technology to unify the color tones. The input data is the acquired video data, and the output is the corrected video data.
[0135] Step 3:
[0136] The server calculates the muscle evaluation score
[0137] The server uses OpenCV to analyze the muscle shading and contours based on the corrected video. Specifically, it uses the Canny edge detection algorithm to perform depth analysis. The input data is the corrected video data, and the output is a muscle evaluation score. The server then calculates the muscle evaluation score using a proprietary algorithm.
[0138] Step 4:
[0139] The server delivers the corrected video and muscle evaluation score in real time.
[0140] The server delivers the corrected video and muscle evaluation scores in real time via WebSocket. The muscle evaluation scores are overlaid on each frame. The input data is the corrected video and muscle evaluation scores, and the output is a stream delivered in real time. Viewers and judges can view this using a browser or a dedicated app.
[0141] Step 5:
[0142] The device receives the video and evaluation data.
[0143] The device receives the corrected video and muscle evaluation data sent from the server via WebSocket. The input data is a real-time stream from the server, and the output is the video and evaluation data available on the device. The device then sends this data to the server as HTML5 <video>Displayed using tags and Canvas technology.
[0144] Step 6:
[0145] The device displays the image
[0146] The device displays the received corrected image on the screen. <video>Video data is inserted into the tag and played back in real time. The input data is the corrected video received, and the output is the video displayed on the screen.
[0147] Step 7:
[0148] The device displays the evaluation data.
[0149] The device displays the muscle evaluation score in real time. Specifically, it analyzes the received evaluation score using JavaScript and performs DOM manipulation to draw the numerical value in a specific area. The input data is the muscle evaluation score, and the output is the evaluation score displayed on the screen.
[0150] Step 8:
[0151] The user selects a video
[0152] The user selects the video of the player they want through the terminal. Specifically, they use a remote control or touch interface to select the player they want from the player list on the screen. The input data is the user's selection information, and the output is a video request to the server.
[0153] Step 9:
[0154] The user watches the video
[0155] The user checks the muscle evaluation score in real time while watching the corrected video. Specifically, the user observes the video and evaluation score displayed on the device screen. The input data is the video and evaluation score displayed on the device, and the output is the user's visual information.
[0156] Step 10:
[0157] Users provide feedback
[0158] Users provide feedback about their viewing experience to the system. Specifically, they use the in-app rating form to enter comments and ratings and then press the submit button. The input data is the user's feedback information, and the output is feedback data sent to the server.
[0159] (Application example 1)
[0160] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0161] In conventional virtual fitting rooms, the user's skin tone is not accurately corrected, which can cause the clothes they try on to look different from how they actually look. Furthermore, there is no adequate mechanism for accurately evaluating the fit of clothes when trying them on in real time. This causes a gap between the actual fitting experience and the virtual fitting experience, reducing the reliability of the experience.
[0162] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0163] In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for overlaying a clothing model on the corrected video, means for evaluating the fit of the clothing model, and means for delivering the fit and corrected video in real time, thereby enabling accurate correction of the user's skin color, displaying the appearance of the clothing being tried on in real time, and also enabling evaluation of the fit.
[0164] "Means for inputting video" refers to a device that has the function of receiving video data from a camera or video capture device and incorporating it into the system.
[0165] The "means for correcting skin color in input video" is a device that uses an artificial intelligence model to perform processing to equalize skin tones in the acquired video and display it accurately.
[0166] The "means for overlaying a clothing model on a corrected image" is a device that performs processing to provide a realistic fitting sensation by overlaying a virtual clothing model on an image with corrected skin color.
[0167] The "means for evaluating the fit of a clothing model" is a device that performs processing to analyze and evaluate the fit of a clothing model in a corrected image.
[0168] The "means for delivering fit and corrected image in real time" refers to a device that includes a communication means and a display means for visually providing the user with the corrected image and the fit of the garment in real time.
[0169] The present invention is a system for correcting a user's skin tone in a virtual fitting room and accurately evaluating the fit of clothing. The following describes the specific processing of each element of this system and the program.
[0170] Server-side processing
[0171] The server performs the process using the following means.
[0172] 1. Acquiring footage
[0173] The server acquires video data from cameras or other video acquisition devices. The video is received in real-time streaming. The camera can be a commonly used webcam or a smartphone camera.
[0174] 2. Skin Tone Correction
[0175] The server analyzes the received video data and uses an artificial intelligence model to correct the players' skin tones. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and processes them to unify their color tones. This process uses deep learning frameworks such as TensorFlow and Keras.
[0176] 3. Overlaying the clothing model
[0177] The server overlays the virtual clothing model onto the corrected image using the OpenCV library, adjusting the position and size of the clothing to fit the user's body shape.
[0178] 4. Fit evaluation
[0179] The server evaluates the fit based on the overlaid clothing model. The fit evaluation algorithm analyzes how well the shape and size of the clothing fits the user's body type and calculates a score. This process also uses deep learning algorithms.
[0180] 5. Real-time streaming
[0181] The server delivers the corrected video and fit evaluation scores in real time using communication protocols such as WebSocket and RTSP.
[0182] Terminal side processing
[0183] The terminal receives the corrected image and fit evaluation data sent from the server and displays them to the user in the most optimal form.
[0184] 1. Receiving video and evaluation data
[0185] The terminal receives the corrected video and fit evaluation data sent from the server.
[0186] 2. Displaying images
[0187] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[0188] 3. Displaying evaluation data
[0189] The device displays the fit evaluation score in real time, which is rendered as a number in a specific area using DOM manipulation.
[0190] User Actions
[0191] Users can perform the following operations through the device:
[0192] 1. Video selection
[0193] Users can use the device to select the image of the desired garment, and then use the remote control or touchscreen interface to select the garment they want to try on.
[0194] 2. Watching the video
[0195] Users can view the corrected footage in real time and check their fit evaluation score, and can even switch to different clothing while viewing.
[0196] 3. Providing Feedback
[0197] Users can provide feedback about their viewing experience to the system, for example by entering and submitting comments and ratings using a rating form within the application.
[0198] Specific examples
[0199] Server operation example: Video is acquired in real time, skin color is corrected using an AI model, and the video is distributed with a clothing model overlaid. The server repeats this process to accumulate fit evaluation data for each garment.
[0200] Example of device operation: The corrected image and fit evaluation score received from the server are displayed on the screen. When the user presses a button to switch to the image of a different garment, a different corrected image and evaluation score are displayed.
[0201] User experience example: Users use a smartphone or tablet application to select and try on clothing items, consider purchasing them based on real-time evaluation data, and submit feedback after the try-on experience using an in-app evaluation form.
[0202] Prompt Sentence Examples
[0203] "I want to develop a virtual fitting room application that captures video in real time, corrects skin color using an AI model, and overlays selected clothing to evaluate fit. Users can capture video of themselves via smartphone or head-mounted display to see how the clothing will look on them. The program uses Python, OpenCV, and Keras."
[0204] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0205] Step 1:
[0206] The server acquires video data in real time from cameras and other video acquisition devices. Specifically, it captures the input signal from the camera and converts it into video data. This process uses the OpenCV library. The input is the video signal from the camera, and the output is the raw video data.
[0207] Step 2:
[0208] The server analyzes the acquired video data using an artificial intelligence model and performs skin color correction. Specifically, it uses a deep learning model (using TensorFlow or Keras) to recognize skin-colored areas in the video and uniformly correct their color tone. The input is RAW video data, and the output is video data with corrected skin color.
[0209] Step 3:
[0210] The server overlays a virtual clothing model on the video data with corrected skin color. Specifically, it performs a process of overlaying an image of the clothing model on the corrected video data. It performs image synthesis using the OpenCV library. The input is the video data with corrected skin color and image data of the clothing model, and the output is video data with the clothing model overlaid.
[0211] Step 4:
[0212] The server evaluates the fit of the clothing based on the overlaid video data. Using a fit evaluation algorithm (deep learning model), it analyzes how well the shape and size of the clothing in the video fits the user's body type and calculates a score. The input is the overlaid video data, and the output is a fit evaluation score.
[0213] Step 5:
[0214] The server delivers the corrected video data and fit evaluation scores in real time. Specifically, it sends the video data and evaluation scores to the user device using a distribution protocol (WebSocket or RTSP). The input is the overlaid video data and fit evaluation scores, and the output is real-time delivery to the user device.
[0215] Step 6:
[0216] The device receives the corrected video and fit evaluation data sent from the server. Specifically, it receives data from the server using WebSocket or RTSP protocols. The input is the video data and evaluation data from the server, and the output is the received data.
[0217] Step 7:
[0218] The device displays the corrected image on the screen. <video>It uses tag and canvas technology to display video in real time. The input is the received video data and the output is the video display to the user.
[0219] Step 8:
[0220] The device displays the fit evaluation score on the screen in real time. Specifically, it uses JavaScript DOM manipulation to draw the evaluation score as a numerical value in a specific area. The input is the fit evaluation data, and the output is the score displayed to the user.
[0221] Step 9:
[0222] The user selects the video of the desired garment using the device and tries it on. The streaming video is controlled using a remote control or touchscreen interface. The input is the user's selection, and the output is a video display of the selected garment.
[0223] Step 10:
[0224] After trying on the products, users provide feedback in an evaluation form within the application. Specifically, they input and submit comments and ratings through the application. The input is the user's feedback, and the output is the transmission of feedback data to the server.
[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0226] This invention integrates a system that corrects the skin color of athletes in bodybuilding competitions and achieves fair muscle evaluation, with an emotion engine that recognizes the user's emotions. The following describes the specific processing of each element of the system and the program. The invention is realized by the server, terminal, and user entities.
[0227] Server-side processing
[0228] 1. Acquiring footage
[0229] The server acquires video data from cameras and other video acquisition devices, and the video data is received in real-time streaming.
[0230] 2. Skin Tone Correction
[0231] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[0232] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[0233] 3. Muscle Assessment
[0234] The corrected image is then further analyzed to detect muscle shadows and contours. The server uses edge detection and shadow analysis algorithms to evaluate muscle prominence.
[0235] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[0236] 4. Real-time streaming
[0237] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[0238] 5. Emotional Data Processing
[0239] The server receives emotional data sent from the user's device, generates feedback in real time based on that data, and sends it to the device.
[0240] Emotion engine integration
[0241] 1. Collecting Emotional Data
[0242] The device analyzes the user's facial expressions using a camera, and the emotion engine recognizes the user's emotions. The emotion engine analyzes the facial expressions using facial feature points in the video and generates emotion data.
[0243] 2. Real-time transmission
[0244] The device transmits the acquired emotion data to a server in real time using a standard protocol.
[0245] 3. Feedback Generation
[0246] The server analyzes the emotion data and generates feedback based on the user's emotion, which may provide, for example, detailed information about a specific player or a cheering message to enhance the user's viewing experience.
[0247] Terminal side processing
[0248] 1. Receiving video and evaluation data
[0249] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[0250] 2. Displaying images
[0251] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[0252] 3. Displaying evaluation data
[0253] The device displays the muscle evaluation score in real time, which is displayed in a specific area using DOM manipulation.
[0254] 4. Acquiring and sending emotion data
[0255] The device uses an emotion engine to analyze the user's facial expressions and transmits the resulting emotion data to the server.
[0256] User Actions
[0257] 1. Video selection
[0258] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[0259] 2. Watching the video
[0260] Users can view the enhanced footage of their selected athletes while viewing their muscle assessment scores, with footage and scores updated in real time.
[0261] 3. Checking emotional feedback
[0262] As users watch, they see real-time emotional feedback sent from the server, often in the form of a pop-up on their screen.
[0263] 4. Providing Feedback
[0264] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[0265] Specific examples
[0266] Example of server operation: Captures video in real time, corrects skin color using an AI model, and delivers the corrected video along with muscle analysis results. Data is also collected from the emotion engine, and feedback is generated based on that.
[0267] Example of device operation: The corrected video and muscle evaluation score received from the server are displayed on the screen, and emotional data is obtained from the user's facial expressions and sent to the server.
[0268] Example of user operation: Users use a smartphone or tablet application to select their favorite players, watch them, and enjoy watching the game while referring to real-time evaluation data and emotional feedback. After watching the game, they also submit feedback using the evaluation form in the app.
[0269] In this way, the present invention is a system that realizes fairer and more transparent bodybuilding competitions by integrating an emotion engine in addition to video correction and muscle evaluation, thereby providing convenience and entertainment not only to athletes but also to viewers.
[0270] The processing flow will be explained below.
[0271] Server-side processing flow
[0272] Step 1:
[0273] The server acquires live video from cameras and video acquisition devices. The video signal is received in streaming format and is captured as video data in real time.
[0274] Step 2:
[0275] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[0276] Step 3:
[0277] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[0278] Step 4:
[0279] The corrected video data is then fed into another AI model to analyze muscle shading and contours. The server uses edge detection and shading analysis algorithms to assess muscle prominence.
[0280] Step 5:
[0281] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[0282] Step 6:
[0283] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[0284] Step 7:
[0285] The server receives emotion data sent from the user's device, and the emotion data is transmitted to the server in real time and collected.
[0286] Step 8:
[0287] The server generates real-time feedback based on the received emotional data, and the feedback content is customized to correspond to the user's emotional state.
[0288] Step 9:
[0289] The server then sends the generated feedback to the user's device, where it is sent in real time and used to improve the viewing experience.
[0290] Processing flow on the terminal side
[0291] Step 1:
[0292] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[0293] Step 2:
[0294] The device analyzes the received data and extracts the video data. <video>It is displayed on the screen using tags and canvases.
[0295] Step 3:
[0296] The device also receives muscle evaluation scores and displays them in a specific area on the screen. The evaluation scores are updated in real time through DOM manipulation.
[0297] Step 4:
[0298] The device uses an emotion engine to analyze the user's facial expressions, processing video data from the camera and identifying the user's facial features.
[0299] Step 5:
[0300] The device recognizes the user's emotions based on the analyzed facial expression data, and the emotion engine uses an AI algorithm to determine emotions such as smile, surprise, or anger in real time.
[0301] Step 6:
[0302] The device transmits the acquired emotion data to a server in real time using a standard protocol.
[0303] Step 7:
[0304] The device receives feedback sent from the server, which is displayed on the user's screen in real time.
[0305] User Operation Flow
[0306] Step 1:
[0307] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[0308] Step 2:
[0309] Users can view the retouched footage of their chosen athlete while viewing their muscle assessment score, which is updated in real time and displayed simultaneously.
[0310] Step 3:
[0311] The user's facial expressions are analyzed in real time using the device's camera. The user does not need to perform any special operations; their natural facial expressions are automatically analyzed by the emotion engine.
[0312] Step 4:
[0313] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[0314] Step 5:
[0315] Users can view real-time feedback sent from the server, which includes detailed player information and supportive messages to enhance the viewing experience.
[0316] In this way, complex information is exchanged between the server, terminals, and users in real time, ensuring a fair and transparent bodybuilding competition.The integration of an emotion engine can make the user's viewing experience more personalized and entertaining.
[0317] Example 2
[0318] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0319] In bodybuilding competitions, differences in the skin color of athletes affect the evaluation of their muscles, making it difficult to make fair evaluations. There is also a lack of systems that can reflect viewers' emotions in real time and improve the viewing experience. The present invention aims to solve these problems and provide a system that achieves fair muscle evaluations and an improved viewing experience.
[0320] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0321] In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for evaluating muscles based on the corrected video, means for distributing the muscle evaluation and the corrected video in real time, means for collecting and analyzing emotional data, and means for generating feedback based on the analyzed emotional data. This enables fair muscle evaluation and improves the viewing experience by providing feedback that reflects the viewer's emotions.
[0322] "Means for inputting video" refers to devices and related technologies for receiving video data in real time from cameras or other video capture devices.
[0323] "Means for correcting skin color in input video" refers to devices and related technologies that use image processing technology and artificial intelligence models to adjust skin color to a standard color for received video data.
[0324] "Means for evaluating muscles based on corrected images" refers to a device and related technology that analyzes video data with corrected skin color, detects muscle shadows and contours, and calculates an evaluation score.
[0325] "Means for delivering muscle evaluation and corrected video in real time" refers to devices and related technologies for transmitting corrected video and muscle evaluation results to viewers in real time via the Internet or other communications networks.
[0326] "Means for collecting and analyzing emotional data" refers to a device and related technology that captures a user's facial expressions and movements with a camera, analyzes them, and detects their emotional state.
[0327] "Means for generating feedback based on analyzed emotional data" refers to a device and related technology for analyzing collected emotional data, creating feedback in real time based on the results, and transmitting the feedback to a viewing terminal.
[0328] This invention integrates a system that corrects the skin color of athletes in bodybuilding competitions and achieves fair muscle evaluation, as well as an emotion engine that recognizes the user's emotions. This invention is realized by the three entities: the server, the terminal, and the user.
[0329] The server acquires video data in real time using a camera or other video acquisition device. This video data is received via the RTSP protocol, for example, using FFmpeg. The received video data is then subjected to skin color correction using an artificial intelligence model (for example, TensorFlow or PyTorch). For example, a prompt such as "Perform skin color correction and apply a unified skin color model" is used as input to the AI model.
[0330] The corrected video data is then analyzed for muscle assessment. This analysis uses image processing algorithms such as Canny edge detection and Sobel filtering to detect muscle shadows and contours. A muscle assessment score is calculated based on the results of this detection. An example prompt is "Perform muscle shadow analysis and calculate an assessment score."
[0331] In addition, the server uses NGINX with the RTMP module and WebSocket to deliver the corrected video and muscle evaluation scores in real time, allowing viewers to receive unbiased video and evaluation data in real time.
[0332] The device receives and displays the corrected video and muscle evaluation score sent from the server in real time. <video>Tags and canvas technology are used to display the evaluation data on the screen using JavaScript.
[0333] The device captures the user's facial expressions and movements with a camera and uses an emotion engine (such as Face++ or Microsoft Azure Face API) to collect the user's emotion data. The collected emotion data is sent from the device to the server in real time via WebSocket. This process allows the server to constantly grasp the user's emotions and analyze them in real time.
[0334] Feedback based on emotion data is generated on the server side. For example, cheering messages and detailed information for specific players can be generated and sent to the user's device. This feedback improves the user's viewing experience. An example of a prompt is "Generate cheering messages based on emotion data."
[0335] In this way, the system of the present invention provides a fairer and more transparent evaluation by integrating an emotion engine in addition to skin color correction and muscle evaluation of the video, thereby improving the viewer experience.
[0336] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0337] Step 1:
[0338] Video Acquisition
[0339] The server receives video data in real time from the camera. The camera is installed at the bodybuilding competition venue and captures video using FFmpeg using the RTSP protocol. The input is video data as an RTSP stream, and the output is a series of video frames that can be processed on the server. This process ensures that the server always has the latest video data.
[0340] Step 2:
[0341] Skin Tone Correction
[0342] The server inputs the acquired video data into an AI model (using TensorFlow or PyTorch) to perform skin color correction. Specifically, it analyzes the color information of each pixel in the image, identifies the skin-colored areas, and performs correction. The input is the pixel data for each video frame, and the output is a video frame with the corrected skin color. As a specific example, the prompt text used is "Perform skin color correction and apply a unified skin color model." This process ensures that skin colors appear consistent even under different lighting conditions.
[0343] Step 3:
[0344] Muscle evaluation
[0345] The server analyzes the corrected video frames using an edge detection algorithm (Canny edge detection or Sobel filter) to identify muscle shading and contours. The input is the video frame with the corrected skin tone, and the output is a muscle evaluation score. Shading analysis is performed using an AI model to evaluate the prominence and shape of specific muscles. As a specific example, the prompt "Perform muscle shading analysis and calculate an evaluation score" is used. This process quantifies the definition of the athlete's muscles.
[0346] Step 4:
[0347] Real-time streaming
[0348] The server delivers the corrected video and muscle evaluation scores to viewers in real time. NGINX and the RTMP module are used to simultaneously send the video data and evaluation scores. The input is the corrected video frame and muscle evaluation score, and the output is the real-time video data and evaluation score via the distribution protocol. This processing allows viewers to receive the latest video and evaluation data in real time.
[0349] Step 5:
[0350] Collecting Emotional Data
[0351] The device captures the user's facial expressions with a camera and generates emotion data using an emotion engine (such as Face++ or Microsoft Azure Face API). The input is video data of the user's face, and the output is data indicating the user's emotional state. This process allows the user's emotions to be analyzed in real time.
[0352] Step 6:
[0353] Real-time transmission
[0354] The device transmits the generated emotion data to the server in real time via WebSocket. The input is the emotion data generated by the engine, and the output is a stream of emotion data to the server. This process allows the server to always grasp the latest emotional state of the user.
[0355] Step 7:
[0356] Feedback Generation
[0357] The server generates feedback based on the user's emotional data. The feedback includes cheering messages and detailed information about the players. The input is the analyzed emotional data, and the output is the generated feedback message. As a concrete example, the prompt "Generate cheering messages based on emotional data" is used. This process allows the user to receive information that improves the viewing experience in real time.
[0358] (Application example 2)
[0359] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0360] In conventional painting processes, the task of evaluating the uniformity and texture of the paint is manual, which can lead to unfairness and inconsistency in evaluations. Furthermore, there is a lack of mechanisms for understanding employees' working environment and emotional state in real time to optimize production efficiency. To solve these problems, a system that integrates automatic evaluation of the paint condition with employee emotional feedback is needed.
[0361] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for evaluating the quality of the paint based on the corrected video, means for distributing the paint quality evaluation and the corrected video in real time, and means for acquiring employee emotions in real time and generating feedback based on the emotion data. This automates the evaluation of paint quality, enabling fair and consistent evaluations, and furthermore, utilizing employee emotion data makes it possible to improve production line efficiency and employee satisfaction.
[0362] "Video" means visual data captured by a camera or other image capture device.
[0363] "Input means" refers to a technique for transmitting images to a server using a device such as a camera.
[0364] "Means for correcting skin tone" refers to a technology that uses an AI model to even out skin tones in video.
[0365] The "means for quality evaluation" is a technology that analyzes the uniformity and texture of the paint from the corrected image and calculates the evaluation results.
[0366] "Means for real-time distribution" refers to a technology for instantly distributing corrected video and quality evaluation results over a network.
[0367] "Means of acquiring emotions in real time" refers to technology that uses devices such as cameras to analyze employees' facial expressions and recognize their emotional state.
[0368] The "means for generating feedback based on emotional data" is a technology that uses recognized emotional data to create feedback for improving production line efficiency and employee satisfaction.
[0369] In the embodiment of the present invention, the following steps are important:
[0370] Server-side processing
[0371] Video Acquisition
[0372] The server acquires video data from cameras and other image acquisition devices attached to the painting robot, and transmits the video data to the server in real time, allowing the painting status to be constantly monitored.
[0373] Skin Tone Correction
[0374] Once the video data is sent to the server, the server uses the generative AI model to correct the skin color (paint color) in the video. Specifically, it uses AI image processing technology to make the color tone of the painted surface uniform.
[0375] Quality assessment
[0376] The corrected image is then further analyzed to evaluate the uniformity and texture of the paint. The server performs this evaluation using edge detection and texture analysis algorithms. The server then quantifies the evaluation results, records them in a database, and distributes them in real time.
[0377] Emotional Data Processing
[0378] The server receives video footage from cameras in the facility and generates emotion data by analyzing employees' facial expressions. The emotion engine uses facial feature points to recognize the employee's emotional state. Based on this data, the server generates feedback to help improve the efficiency of the production line and sends it to the terminal.
[0379] Terminal side processing
[0380] Video display
[0381] The terminal receives the corrected video and quality evaluation data sent from the server in real time and displays them on the screen. <video>Tag and canvas technology is used.
[0382] Viewing evaluation data
[0383] The terminal displays the evaluation data on the screen in real time, providing the user with evaluation results on the uniformity and texture of the paint.
[0384] Acquiring and sending emotion data
[0385] The device also uses a built-in camera to analyze the employee's facial expressions and generate emotional data, which is then sent to a server in real time.
[0386] User Actions
[0387] Users can operate the system using a management terminal, check the video footage and evaluation results in real time, and send feedback to the system as needed.
[0388] Specific use cases
[0389] For example, when a new painting process is introduced, a camera attached to the robot captures the painting process and sends the footage to a server. The server then uses an AI model to correct the paint color and evaluate the paint's uniformity and texture. The results are sent to the manager's device in real time. The camera also captures facial expression data from employees, and their emotional state is analyzed. Based on the analysis results, feedback to improve the efficiency of the production line is automatically provided.
[0390] Example prompt sentences to use
[0391] Here is an example of a prompt to input to a generative AI model:
[0392] Analyze employee emotions using the image below and output the results in JSON format. The facial expression data contains specific emotions including smile, anger, sadness, satisfaction, etc.
[0393] image: <base64-encoded-image-data>
[0394] Analyze the paint uniformity and texture for the following real-time video frames, quantify the evaluation results, and output them in JSON format.
[0395] Video Frame: <base64-encoded-frame-data>
[0396] As described above, the present invention provides a system that integrates the painting process and employee emotion recognition, enabling the evaluation of painting quality and the improvement of production efficiency.
[0397] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0398] Step 1:
[0399] The server acquires video data in real time from a camera attached to the painting robot. It receives the video stream sent from the camera and stores it in a buffer for processing. The input is the video data, and the output is the video frames in the buffer. Specifically, it acquires video using the camera's IP address and connection protocol (e.g., RTSP), and processes each frame using a library such as OpenCV.
[0400] Step 2:
[0401] The server inputs the captured video frame into a generative AI model to correct the paint color. The AI model analyzes each pixel of the video frame and evens out the color. The input is the video frame, and the output is a color-corrected frame. Specifically, it calls the AI model, performs inference processing, and applies the color correction algorithm.
[0402] Step 3:
[0403] The server analyzes the color-corrected video frames to evaluate the uniformity and texture of the paint. Edge detection and texture analysis algorithms are used to quantify the evaluation of the paint surface. The input is the color-corrected video frames, and the output is the numerical evaluation results. Specifically, edge detection and texture analysis are performed using OpenCV and an image analysis library, and an evaluation score is calculated.
[0404] Step 4:
[0405] The server delivers the evaluation results and corrected video frames to the administrator's terminal in real time. Real-time data transmission is performed using WebSocket or HTTP streaming. The input is the evaluation results and color-corrected video frames, and the output is the evaluation data and video displayed on the administrator's terminal. Specifically, a WebSocket server is set up and a real-time data stream is transmitted.
[0406] Step 5:
[0407] The server acquires facial expression data of employees from cameras within the facility and generates emotion data using an emotion engine. The emotion engine analyzes facial feature points and recognizes emotional states. The input is the employee's facial expression data, and the output is emotion data. Specifically, it detects faces and analyzes facial expressions to generate emotion data.
[0408] Step 6:
[0409] The server generates feedback based on the emotion data to improve the efficiency of the production line and sends it to the terminal. The input is emotion data, and the output is a feedback message. Specifically, it generates support messages and improvement suggestions based on the analysis results of the emotion data and sends them to the terminal via an HTTP request.
[0410] Step 7:
[0411] The terminal receives the corrected video and evaluation data sent from the server and displays them on the screen. The input is the evaluation data and the corrected video frame, and the output is the information displayed on the terminal's display. The specific operation is as follows: <video>Video is played using tags and canvas technology, and evaluation data is displayed using DOM manipulation.
[0412] Step 8:
[0413] The device uses a built-in camera to analyze the employee's facial expressions, generate emotional data, and send it to the server. The input is the employee's facial expression data, and the output is emotional data sent to the server. Specifically, the device acquires camera images, generates emotional data using facial expression analysis technology, and sends it to the server in real time.
[0414] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0415] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0416] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0417] [Second embodiment]
[0418] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0419] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0420] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0421] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0422] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0423] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0424] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0425] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0426] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0427] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0428] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0429] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0430] The present invention is a system for correcting the skin color of athletes in bodybuilding competitions and achieving fair muscle evaluation. The following describes the specific processing of each system element and program. The present invention is realized by the server, terminals, and users.
[0431] Server-side processing
[0432] 1. Acquiring footage
[0433] The server acquires video data from cameras and other video acquisition devices. The video is received in real-time streaming.
[0434] 2. Skin Tone Correction
[0435] The server analyzes the received video data and uses an AI model to correct the players' skin tones. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and processes them to unify their color tones.
[0436] 3. Muscle Assessment
[0437] The corrected image is further analyzed to detect muscle shadows and contours. The server then calculates a muscle evaluation score based on this. The muscle evaluation algorithm uses edge detection and depth analysis to evaluate the degree to which muscles stand out in detail.
[0438] 4. Real-time streaming
[0439] The server distributes the corrected video and muscle evaluation scores to viewers and judges in real time via the Internet or dedicated communication protocols (e.g., RTSP, WebSocket).
[0440] Terminal side processing
[0441] 1. Receiving video and evaluation data
[0442] The terminal receives the corrected video and muscle evaluation data transmitted from the server.
[0443] 2. Displaying images
[0444] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[0445] 3. Displaying evaluation data
[0446] The device displays the muscle evaluation score in real time, which is rendered as a number in a specific area using DOM manipulation.
[0447] User Actions
[0448] 1. Video selection
[0449] Users can select the video of the player they want to evaluate using the device, for example, a remote control or touch interface.
[0450] 2. Watching the video
[0451] Users can view the corrected footage and check the muscle evaluation score in real time. They can also switch to other athletes while viewing.
[0452] 3. Providing Feedback
[0453] Users can provide feedback about their viewing experience to the system, for example by entering and submitting comments and ratings using a rating form within the application.
[0454] Specific examples
[0455] Example of server operation: Video is acquired in real time, skin color is corrected using an AI model, and the corrected video is distributed along with muscle analysis results. The server repeats this process to accumulate evaluation data for each athlete.
[0456] Example of device operation: The corrected video and muscle evaluation score received from the server are displayed on the screen. When the user presses a button to switch to the video of another athlete, a different corrected video and evaluation score are displayed.
[0457] Example of user operation: Users use the application on their smartphone or tablet to select their favorite players to watch, and enjoy watching the game while referring to real-time evaluation data. After watching the game, they also submit feedback using the evaluation form in the app.
[0458] In this way, the present invention is a system that realizes a fair and transparent bodybuilding competition through video correction and muscle evaluation, thereby providing convenience and entertainment not only to athletes but also to viewers.
[0459] The processing flow will be explained below.
[0460] Server-side processing flow
[0461] Step 1:
[0462] The server acquires live video from cameras and video acquisition devices. The video signal is received in streaming format and is captured as video data in real time.
[0463] Step 2:
[0464] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[0465] Step 3:
[0466] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[0467] Step 4:
[0468] The corrected video data is then fed into another AI model to analyze muscle shading and contours. The server uses edge detection and shading analysis algorithms to assess muscle prominence.
[0469] Step 5:
[0470] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[0471] Step 6:
[0472] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[0473] Processing flow on the terminal side
[0474] Step 1:
[0475] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[0476] Step 2:
[0477] The device analyzes the received data and extracts the video data. <video>It is displayed on the screen using tags and canvases.
[0478] Step 3:
[0479] The device also receives muscle evaluation scores and displays them in a specific area on the screen. The evaluation scores are updated in real time through DOM manipulation.
[0480] User Operation Flow
[0481] Step 1:
[0482] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[0483] Step 2:
[0484] Users can view the retouched footage of their chosen athlete while viewing their muscle assessment score, which is updated in real time and displayed simultaneously.
[0485] Step 3:
[0486] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[0487] In this way, the server, terminals, and users each play their respective roles, and a system is operated that realizes a fair and transparent bodybuilding competition in real time.
[0488] Example 1
[0489] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0490] In existing bodybuilding competitions, the skin color of athletes affects the evaluation, making it difficult to provide a fair muscle evaluation. Additionally, there is a lack of systems that allow judges and viewers to provide fair evaluations in real time. Therefore, an effective method to correct athletes' skin color and provide a fair muscle evaluation is needed.
[0491] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0492] In this invention, the server includes means for acquiring video from a video acquisition device, means for correcting skin color in the acquired video, means for calculating a muscle evaluation score based on the corrected video, means for distributing the corrected video and muscle evaluation score in real time, means for displaying the received corrected video, and means for displaying the muscle evaluation score in real time. This allows for fair evaluation by correcting the skin color of the athlete, and makes it possible to provide fair muscle evaluations in real time to judges and viewers.
[0493] "Video capture device" refers to a camera or other video capture mechanism that captures video data in real time.
[0494] The "skin color correction means" is a technology that performs processing to ensure consistent skin color tones of players in the captured video.
[0495] The "means for calculating muscle evaluation scores" is a technology that analyzes the corrected video to evaluate the shading and contours of the player's muscles and convert them into a numerical score.
[0496] "Real-time distribution means" means technology that distributes the corrected footage and muscle evaluation scores to viewers and judges in real time without delay via the Internet or other communications protocols.
[0497] The "means for displaying corrected images" is a technique for displaying corrected images received from the server on the screen of the terminal.
[0498] "Means for displaying muscle evaluation scores in real time" refers to a technology that instantly displays muscle evaluation scores received from a server on the screen of a terminal.
[0499] A "generative AI model" is an artificial intelligence model that learns from large amounts of data and is designed to perform specific tasks.
[0500] The present invention is a system for correcting the skin color of athletes in bodybuilding competitions and achieving fair muscle evaluation. The present invention is realized by the server, terminals, and users.
[0501] Server-side processing
[0502] The server first acquires video data in real time from a video capture device, such as a high-resolution camera. The video is then sent to the server using the RTSP protocol, where it is received and processed by software such as FFmpeg.
[0503] The server then uses generative AI models such as TensorFlow to correct the skin tone of the captured video. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and uses techniques such as histogram matching to unify the skin tone.
[0504] The corrected video is analyzed for muscle shading and contours using OpenCV. The server performs this analysis using the Canny edge detection algorithm and depth analysis technology, and then calculates a muscle evaluation score using a proprietary algorithm based on the obtained data.
[0505] The server also uses WebSocket to deliver the corrected video and muscle evaluation scores to viewers and judges in real time. The muscle evaluation scores are overlaid on each frame, and viewers can view them in a browser or a dedicated app.
[0506] Terminal side processing
[0507] The device receives the corrected video and muscle evaluation data sent from the server using WebSocket as the reception protocol.
[0508] The device analyzes the received data and <video>The corrected image is displayed in real time using tags and Canvas technology, and the muscle evaluation score is displayed on the screen using JavaScript and DOM manipulation, allowing users to instantly check the corrected image and real-time muscle evaluation score.
[0509] User operations
[0510] The user selects the video of the desired player through the device. Using a remote control or touch interface, the desired player can be selected from a list of players on the screen. The video of the selected player is requested from the server, and the corrected video is sent to the device.
[0511] Users can check their muscle evaluation score in real time while watching the corrected video. After watching, they can provide feedback about their viewing experience using the in-app rating form. This feedback data is sent to the server and used for future analysis and improvements.
[0512] Specific examples
[0513] Example of server operation: The server acquires images in real time from a high-resolution camera and performs skin color correction using TensorFlow. Next, it analyzes muscle shading and contours using OpenCV, and delivers the corrected images and evaluation data in real time via WebSocket.
[0514] Example prompt sentence:
[0515] Capture video in real time, correct skin color using TensorFlow, analyze muscles using Canny edge detection, and deliver the results via WebSocket.
[0516] Example of device operation: The device receives the corrected video and muscle evaluation score from the server and sends them to the HTML5 <video>The tag and Canvas are used to display the player on the screen. When the user presses a button on the remote control to switch players, the new video and evaluation score are displayed.
[0517] Example prompt sentence:
[0518] Video received from the server is processed as HTML5 <video>Play with tags and use Canvas to display muscle evaluation scores. Switch players when the user selects them with the remote.
[0519] Example of user operation: The user uses the smartphone app to select their favorite player by touch operation and watch. After watching, they submit feedback using the evaluation form in the app.
[0520] Example prompt sentence:
[0521] Select a player by touching the smartphone app, and view the adjusted footage and muscle evaluation score. After viewing, submit your feedback in the evaluation form.
[0522] Thus, the present invention is a system that realizes fair and transparent bodybuilding competitions through video correction and muscle evaluation, and provides convenience and entertainment for both athletes and viewers.
[0523] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0524] Step 1:
[0525] The server acquires the video
[0526] Specifically, the server captures real-time video of the players using a high-resolution camera. This video data is sent to the server via the RTSP protocol. The input data is the raw video stream, and the output is the video data available within the server. The server receives and stores the video data using software such as FFmpeg.
[0527] Step 2:
[0528] The server corrects skin tones
[0529] After acquiring the video data, the server uses a generative AI model such as TensorFlow to perform skin tone correction. Specifically, the AI model receives the video data as input, recognizes skin-tone areas, and uses histogram matching technology to unify the color tones. The input data is the acquired video data, and the output is the corrected video data.
[0530] Step 3:
[0531] The server calculates the muscle evaluation score
[0532] The server uses OpenCV to analyze the muscle shading and contours based on the corrected video. Specifically, it uses the Canny edge detection algorithm to perform depth analysis. The input data is the corrected video data, and the output is a muscle evaluation score. The server then calculates the muscle evaluation score using a proprietary algorithm.
[0533] Step 4:
[0534] The server delivers the corrected video and muscle evaluation score in real time.
[0535] The server delivers the corrected video and muscle evaluation scores in real time via WebSocket. The muscle evaluation scores are overlaid on each frame. The input data is the corrected video and muscle evaluation scores, and the output is a stream delivered in real time. Viewers and judges can view this using a browser or a dedicated app.
[0536] Step 5:
[0537] The device receives the video and evaluation data.
[0538] The device receives the corrected video and muscle evaluation data sent from the server via WebSocket. The input data is a real-time stream from the server, and the output is the video and evaluation data available on the device. The device then sends this data to the server as HTML5 <video>Displayed using tags and Canvas technology.
[0539] Step 6:
[0540] The device displays the image
[0541] The device displays the received corrected image on the screen. <video>Video data is inserted into the tag and played back in real time. The input data is the corrected video received, and the output is the video displayed on the screen.
[0542] Step 7:
[0543] The device displays the evaluation data.
[0544] The device displays the muscle evaluation score in real time. Specifically, it analyzes the received evaluation score using JavaScript and performs DOM manipulation to draw the numerical value in a specific area. The input data is the muscle evaluation score, and the output is the evaluation score displayed on the screen.
[0545] Step 8:
[0546] The user selects a video
[0547] The user selects the video of the player they want through the terminal. Specifically, they use a remote control or touch interface to select the player they want from the player list on the screen. The input data is the user's selection information, and the output is a video request to the server.
[0548] Step 9:
[0549] The user watches the video
[0550] The user checks the muscle evaluation score in real time while watching the corrected video. Specifically, the user observes the video and evaluation score displayed on the device screen. The input data is the video and evaluation score displayed on the device, and the output is the user's visual information.
[0551] Step 10:
[0552] Users provide feedback
[0553] Users provide feedback about their viewing experience to the system. Specifically, they use the in-app rating form to enter comments and ratings and then press the submit button. The input data is the user's feedback information, and the output is feedback data sent to the server.
[0554] (Application example 1)
[0555] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0556] In conventional virtual fitting rooms, the user's skin tone is not accurately corrected, which can cause the clothes they try on to look different from how they actually look. Furthermore, there is no adequate mechanism for accurately evaluating the fit of clothes when trying them on in real time. This causes a gap between the actual fitting experience and the virtual fitting experience, reducing the reliability of the experience.
[0557] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0558] In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for overlaying a clothing model on the corrected video, means for evaluating the fit of the clothing model, and means for delivering the fit and corrected video in real time, thereby enabling accurate correction of the user's skin color, displaying the appearance of the clothing being tried on in real time, and also enabling evaluation of the fit.
[0559] "Means for inputting video" refers to a device that has the function of receiving video data from a camera or video capture device and incorporating it into the system.
[0560] The "means for correcting skin color in input video" is a device that uses an artificial intelligence model to perform processing to equalize skin tones in the acquired video and display it accurately.
[0561] The "means for overlaying a clothing model on a corrected image" is a device that performs processing to provide a realistic fitting sensation by overlaying a virtual clothing model on an image with corrected skin color.
[0562] The "means for evaluating the fit of a clothing model" is a device that performs processing to analyze and evaluate the fit of a clothing model in a corrected image.
[0563] The "means for delivering fit and corrected image in real time" refers to a device that includes a communication means and a display means for visually providing the user with the corrected image and the fit of the garment in real time.
[0564] The present invention is a system for correcting a user's skin tone in a virtual fitting room and accurately evaluating the fit of clothing. The following describes the specific processing of each element of this system and the program.
[0565] Server-side processing
[0566] The server performs the process using the following means.
[0567] 1. Acquiring footage
[0568] The server acquires video data from cameras or other video acquisition devices. The video is received in real-time streaming. The camera can be a commonly used webcam or a smartphone camera.
[0569] 2. Skin Tone Correction
[0570] The server analyzes the received video data and uses an artificial intelligence model to correct the players' skin tones. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and processes them to unify their color tones. This process uses deep learning frameworks such as TensorFlow and Keras.
[0571] 3. Overlaying the clothing model
[0572] The server overlays the virtual clothing model onto the corrected image using the OpenCV library, adjusting the position and size of the clothing to fit the user's body shape.
[0573] 4. Fit evaluation
[0574] The server evaluates the fit based on the overlaid clothing model. The fit evaluation algorithm analyzes how well the shape and size of the clothing fits the user's body type and calculates a score. This process also uses deep learning algorithms.
[0575] 5. Real-time streaming
[0576] The server delivers the corrected video and fit evaluation scores in real time using communication protocols such as WebSocket and RTSP.
[0577] Terminal side processing
[0578] The terminal receives the corrected image and fit evaluation data sent from the server and displays them to the user in the most optimal form.
[0579] 1. Receiving video and evaluation data
[0580] The terminal receives the corrected video and fit evaluation data sent from the server.
[0581] 2. Displaying images
[0582] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[0583] 3. Displaying evaluation data
[0584] The device displays the fit evaluation score in real time, which is rendered as a number in a specific area using DOM manipulation.
[0585] User Actions
[0586] Users can perform the following operations through the device:
[0587] 1. Video selection
[0588] Users can use the device to select the image of the desired garment, and then use the remote control or touchscreen interface to select the garment they want to try on.
[0589] 2. Watching the video
[0590] Users can view the corrected footage in real time and check their fit evaluation score, and can even switch to different clothing while viewing.
[0591] 3. Providing Feedback
[0592] Users can provide feedback about their viewing experience to the system, for example by entering and submitting comments and ratings using a rating form within the application.
[0593] Specific examples
[0594] Server operation example: Video is acquired in real time, skin color is corrected using an AI model, and the video is distributed with a clothing model overlaid. The server repeats this process to accumulate fit evaluation data for each garment.
[0595] Example of device operation: The corrected image and fit evaluation score received from the server are displayed on the screen. When the user presses a button to switch to the image of a different garment, a different corrected image and evaluation score are displayed.
[0596] User experience example: Users use a smartphone or tablet application to select and try on clothing items, consider purchasing them based on real-time evaluation data, and submit feedback after the try-on experience using an in-app evaluation form.
[0597] Prompt Sentence Examples
[0598] "I want to develop a virtual fitting room application that captures video in real time, corrects skin color using an AI model, and overlays selected clothing to evaluate fit. Users can capture video of themselves via smartphone or head-mounted display to see how the clothing will look on them. The program uses Python, OpenCV, and Keras."
[0599] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0600] Step 1:
[0601] The server acquires video data in real time from cameras and other video acquisition devices. Specifically, it captures the input signal from the camera and converts it into video data. This process uses the OpenCV library. The input is the video signal from the camera, and the output is the raw video data.
[0602] Step 2:
[0603] The server analyzes the acquired video data using an artificial intelligence model and performs skin color correction. Specifically, it uses a deep learning model (using TensorFlow or Keras) to recognize skin-colored areas in the video and uniformly correct their color tone. The input is RAW video data, and the output is video data with corrected skin color.
[0604] Step 3:
[0605] The server overlays a virtual clothing model on the video data with corrected skin color. Specifically, it performs a process of overlaying an image of the clothing model on the corrected video data. It performs image synthesis using the OpenCV library. The input is the video data with corrected skin color and image data of the clothing model, and the output is video data with the clothing model overlaid.
[0606] Step 4:
[0607] The server evaluates the fit of the clothing based on the overlaid video data. Using a fit evaluation algorithm (deep learning model), it analyzes how well the shape and size of the clothing in the video fits the user's body type and calculates a score. The input is the overlaid video data, and the output is a fit evaluation score.
[0608] Step 5:
[0609] The server delivers the corrected video data and fit evaluation scores in real time. Specifically, it sends the video data and evaluation scores to the user device using a distribution protocol (WebSocket or RTSP). The input is the overlaid video data and fit evaluation scores, and the output is real-time delivery to the user device.
[0610] Step 6:
[0611] The device receives the corrected video and fit evaluation data sent from the server. Specifically, it receives data from the server using WebSocket or RTSP protocols. The input is the video data and evaluation data from the server, and the output is the received data.
[0612] Step 7:
[0613] The device displays the corrected image on the screen. <video>It uses tag and canvas technology to display video in real time. The input is the received video data and the output is the video display to the user.
[0614] Step 8:
[0615] The device displays the fit evaluation score on the screen in real time. Specifically, it uses JavaScript DOM manipulation to draw the evaluation score as a numerical value in a specific area. The input is the fit evaluation data, and the output is the score displayed to the user.
[0616] Step 9:
[0617] The user selects the video of the desired garment using the device and tries it on. The streaming video is controlled using a remote control or touchscreen interface. The input is the user's selection, and the output is a video display of the selected garment.
[0618] Step 10:
[0619] After trying on the products, users provide feedback in an evaluation form within the application. Specifically, they input and submit comments and ratings through the application. The input is the user's feedback, and the output is the transmission of feedback data to the server.
[0620] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0621] This invention integrates a system that corrects the skin color of athletes in bodybuilding competitions and achieves fair muscle evaluation, with an emotion engine that recognizes the user's emotions. The following describes the specific processing of each element of the system and the program. The invention is realized by the server, terminal, and user entities.
[0622] Server-side processing
[0623] 1. Acquiring footage
[0624] The server acquires video data from cameras and other video acquisition devices, and the video data is received in real-time streaming.
[0625] 2. Skin Tone Correction
[0626] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[0627] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[0628] 3. Muscle Assessment
[0629] The corrected image is then further analyzed to detect muscle shadows and contours. The server uses edge detection and shadow analysis algorithms to evaluate muscle prominence.
[0630] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[0631] 4. Real-time streaming
[0632] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[0633] 5. Emotional Data Processing
[0634] The server receives emotional data sent from the user's device, generates feedback in real time based on that data, and sends it to the device.
[0635] Emotion engine integration
[0636] 1. Collecting Emotional Data
[0637] The device analyzes the user's facial expressions using a camera, and the emotion engine recognizes the user's emotions. The emotion engine analyzes the facial expressions using facial feature points in the video and generates emotion data.
[0638] 2. Real-time transmission
[0639] The device transmits the acquired emotion data to a server in real time using a standard protocol.
[0640] 3. Feedback Generation
[0641] The server analyzes the emotion data and generates feedback based on the user's emotion, which may provide, for example, detailed information about a specific player or a cheering message to enhance the user's viewing experience.
[0642] Terminal side processing
[0643] 1. Receiving video and evaluation data
[0644] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[0645] 2. Displaying images
[0646] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[0647] 3. Displaying evaluation data
[0648] The device displays the muscle evaluation score in real time, which is displayed in a specific area using DOM manipulation.
[0649] 4. Acquiring and sending emotion data
[0650] The device uses an emotion engine to analyze the user's facial expressions and transmits the resulting emotion data to the server.
[0651] User Actions
[0652] 1. Video selection
[0653] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[0654] 2. Watching the video
[0655] Users can view the enhanced footage of their selected athletes while viewing their muscle assessment scores, with footage and scores updated in real time.
[0656] 3. Checking emotional feedback
[0657] As users watch, they see real-time emotional feedback sent from the server, often in the form of a pop-up on their screen.
[0658] 4. Providing Feedback
[0659] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[0660] Specific examples
[0661] Example of server operation: Captures video in real time, corrects skin color using an AI model, and delivers the corrected video along with muscle analysis results. Data is also collected from the emotion engine, and feedback is generated based on that.
[0662] Example of device operation: The corrected video and muscle evaluation score received from the server are displayed on the screen, and emotional data is obtained from the user's facial expressions and sent to the server.
[0663] Example of user operation: Users use a smartphone or tablet application to select their favorite players, watch them, and enjoy watching the game while referring to real-time evaluation data and emotional feedback. After watching the game, they also submit feedback using the evaluation form in the app.
[0664] In this way, the present invention is a system that realizes fairer and more transparent bodybuilding competitions by integrating an emotion engine in addition to video correction and muscle evaluation, thereby providing convenience and entertainment not only to athletes but also to viewers.
[0665] The processing flow will be explained below.
[0666] Server-side processing flow
[0667] Step 1:
[0668] The server acquires live video from cameras and video acquisition devices. The video signal is received in streaming format and is captured as video data in real time.
[0669] Step 2:
[0670] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[0671] Step 3:
[0672] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[0673] Step 4:
[0674] The corrected video data is then fed into another AI model to analyze muscle shading and contours. The server uses edge detection and shading analysis algorithms to assess muscle prominence.
[0675] Step 5:
[0676] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[0677] Step 6:
[0678] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[0679] Step 7:
[0680] The server receives emotion data sent from the user's device, and the emotion data is transmitted to the server in real time and collected.
[0681] Step 8:
[0682] The server generates real-time feedback based on the received emotional data, and the feedback content is customized to correspond to the user's emotional state.
[0683] Step 9:
[0684] The server then sends the generated feedback to the user's device, where it is sent in real time and used to improve the viewing experience.
[0685] Processing flow on the terminal side
[0686] Step 1:
[0687] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[0688] Step 2:
[0689] The device analyzes the received data and extracts the video data. <video>It is displayed on the screen using tags and canvases.
[0690] Step 3:
[0691] The device also receives muscle evaluation scores and displays them in a specific area on the screen. The evaluation scores are updated in real time through DOM manipulation.
[0692] Step 4:
[0693] The device uses an emotion engine to analyze the user's facial expressions, processing video data from the camera and identifying the user's facial features.
[0694] Step 5:
[0695] The device recognizes the user's emotions based on the analyzed facial expression data, and the emotion engine uses an AI algorithm to determine emotions such as smile, surprise, or anger in real time.
[0696] Step 6:
[0697] The device transmits the acquired emotion data to a server in real time using a standard protocol.
[0698] Step 7:
[0699] The device receives feedback sent from the server, which is displayed on the user's screen in real time.
[0700] User Operation Flow
[0701] Step 1:
[0702] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[0703] Step 2:
[0704] Users can view the retouched footage of their chosen athlete while viewing their muscle assessment score, which is updated in real time and displayed simultaneously.
[0705] Step 3:
[0706] The user's facial expressions are analyzed in real time using the device's camera. The user does not need to perform any special operations; their natural facial expressions are automatically analyzed by the emotion engine.
[0707] Step 4:
[0708] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[0709] Step 5:
[0710] Users can view real-time feedback sent from the server, which includes detailed player information and supportive messages to enhance the viewing experience.
[0711] In this way, complex information is exchanged between the server, terminals, and users in real time, ensuring a fair and transparent bodybuilding competition.The integration of an emotion engine can make the user's viewing experience more personalized and entertaining.
[0712] Example 2
[0713] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0714] In bodybuilding competitions, differences in the skin color of athletes affect the evaluation of their muscles, making it difficult to make fair evaluations. There is also a lack of systems that can reflect viewers' emotions in real time and improve the viewing experience. The present invention aims to solve these problems and provide a system that achieves fair muscle evaluations and an improved viewing experience.
[0715] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0716] In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for evaluating muscles based on the corrected video, means for distributing the muscle evaluation and the corrected video in real time, means for collecting and analyzing emotional data, and means for generating feedback based on the analyzed emotional data. This enables fair muscle evaluation and improves the viewing experience by providing feedback that reflects the viewer's emotions.
[0717] "Means for inputting video" refers to devices and related technologies for receiving video data in real time from cameras or other video capture devices.
[0718] "Means for correcting skin color in input video" refers to devices and related technologies that use image processing technology and artificial intelligence models to adjust skin color to a standard color for received video data.
[0719] "Means for evaluating muscles based on corrected images" refers to a device and related technology that analyzes video data with corrected skin color, detects muscle shadows and contours, and calculates an evaluation score.
[0720] "Means for delivering muscle evaluation and corrected video in real time" refers to devices and related technologies for transmitting corrected video and muscle evaluation results to viewers in real time via the Internet or other communications networks.
[0721] "Means for collecting and analyzing emotional data" refers to a device and related technology that captures a user's facial expressions and movements with a camera, analyzes them, and detects their emotional state.
[0722] "Means for generating feedback based on analyzed emotional data" refers to a device and related technology for analyzing collected emotional data, creating feedback in real time based on the results, and transmitting the feedback to a viewing terminal.
[0723] This invention integrates a system that corrects the skin color of athletes in bodybuilding competitions and achieves fair muscle evaluation, as well as an emotion engine that recognizes the user's emotions. This invention is realized by the three entities: the server, the terminal, and the user.
[0724] The server acquires video data in real time using a camera or other video acquisition device. This video data is received via the RTSP protocol, for example, using FFmpeg. The received video data is then subjected to skin color correction using an artificial intelligence model (for example, TensorFlow or PyTorch). For example, a prompt such as "Perform skin color correction and apply a unified skin color model" is used as input to the AI model.
[0725] The corrected video data is then analyzed for muscle assessment. This analysis uses image processing algorithms such as Canny edge detection and Sobel filtering to detect muscle shadows and contours. A muscle assessment score is calculated based on the results of this detection. An example prompt is "Perform muscle shadow analysis and calculate an assessment score."
[0726] In addition, the server uses NGINX with the RTMP module and WebSocket to deliver the corrected video and muscle evaluation scores in real time, allowing viewers to receive unbiased video and evaluation data in real time.
[0727] The device receives and displays the corrected video and muscle evaluation score sent from the server in real time. <video>Tags and canvas technology are used to display the evaluation data on the screen using JavaScript.
[0728] The device captures the user's facial expressions and movements with a camera and uses an emotion engine (such as Face++ or Microsoft Azure Face API) to collect the user's emotion data. The collected emotion data is sent from the device to the server in real time via WebSocket. This process allows the server to constantly grasp the user's emotions and analyze them in real time.
[0729] Feedback based on emotion data is generated on the server side. For example, cheering messages and detailed information for specific players can be generated and sent to the user's device. This feedback improves the user's viewing experience. An example of a prompt is "Generate cheering messages based on emotion data."
[0730] In this way, the system of the present invention provides a fairer and more transparent evaluation by integrating an emotion engine in addition to skin color correction and muscle evaluation of the video, thereby improving the viewer experience.
[0731] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0732] Step 1:
[0733] Video Acquisition
[0734] The server receives video data in real time from the camera. The camera is installed at the bodybuilding competition venue and captures video using FFmpeg using the RTSP protocol. The input is video data as an RTSP stream, and the output is a series of video frames that can be processed on the server. This process ensures that the server always has the latest video data.
[0735] Step 2:
[0736] Skin Tone Correction
[0737] The server inputs the acquired video data into an AI model (using TensorFlow or PyTorch) to perform skin color correction. Specifically, it analyzes the color information of each pixel in the image, identifies the skin-colored areas, and performs correction. The input is the pixel data for each video frame, and the output is a video frame with the corrected skin color. As a specific example, the prompt text used is "Perform skin color correction and apply a unified skin color model." This process ensures that skin colors appear consistent even under different lighting conditions.
[0738] Step 3:
[0739] Muscle evaluation
[0740] The server analyzes the corrected video frames using an edge detection algorithm (Canny edge detection or Sobel filter) to identify muscle shading and contours. The input is the video frame with the corrected skin tone, and the output is a muscle evaluation score. Shading analysis is performed using an AI model to evaluate the prominence and shape of specific muscles. As a specific example, the prompt "Perform muscle shading analysis and calculate an evaluation score" is used. This process quantifies the definition of the athlete's muscles.
[0741] Step 4:
[0742] Real-time streaming
[0743] The server delivers the corrected video and muscle evaluation scores to viewers in real time. NGINX and the RTMP module are used to simultaneously send the video data and evaluation scores. The input is the corrected video frame and muscle evaluation score, and the output is the real-time video data and evaluation score via the distribution protocol. This processing allows viewers to receive the latest video and evaluation data in real time.
[0744] Step 5:
[0745] Collecting Emotional Data
[0746] The device captures the user's facial expressions with a camera and generates emotion data using an emotion engine (such as Face++ or Microsoft Azure Face API). The input is video data of the user's face, and the output is data indicating the user's emotional state. This process allows the user's emotions to be analyzed in real time.
[0747] Step 6:
[0748] Real-time transmission
[0749] The device transmits the generated emotion data to the server in real time via WebSocket. The input is the emotion data generated by the engine, and the output is a stream of emotion data to the server. This process allows the server to always grasp the latest emotional state of the user.
[0750] Step 7:
[0751] Feedback Generation
[0752] The server generates feedback based on the user's emotional data. The feedback includes cheering messages and detailed information about the players. The input is the analyzed emotional data, and the output is the generated feedback message. As a concrete example, the prompt "Generate cheering messages based on emotional data" is used. This process allows the user to receive information that improves the viewing experience in real time.
[0753] (Application example 2)
[0754] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0755] In conventional painting processes, the task of evaluating the uniformity and texture of the paint is manual, which can lead to unfairness and inconsistency in evaluations. Furthermore, there is a lack of mechanisms for understanding employees' working environment and emotional state in real time to optimize production efficiency. To solve these problems, a system that integrates automatic evaluation of the paint condition with employee emotional feedback is needed.
[0756] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for evaluating the quality of the paint based on the corrected video, means for distributing the paint quality evaluation and the corrected video in real time, and means for acquiring employee emotions in real time and generating feedback based on the emotion data. This automates the evaluation of paint quality, enabling fair and consistent evaluations, and furthermore, utilizing employee emotion data makes it possible to improve production line efficiency and employee satisfaction.
[0757] "Video" means visual data captured by a camera or other image capture device.
[0758] "Input means" refers to a technique for transmitting images to a server using a device such as a camera.
[0759] "Means for correcting skin tone" refers to a technology that uses an AI model to even out skin tones in video.
[0760] The "means for quality evaluation" is a technology that analyzes the uniformity and texture of the paint from the corrected image and calculates the evaluation results.
[0761] "Means for real-time distribution" refers to a technology for instantly distributing corrected video and quality evaluation results over a network.
[0762] "Means of acquiring emotions in real time" refers to technology that uses devices such as cameras to analyze employees' facial expressions and recognize their emotional state.
[0763] The "means for generating feedback based on emotional data" is a technology that uses recognized emotional data to create feedback for improving production line efficiency and employee satisfaction.
[0764] In the embodiment of the present invention, the following steps are important:
[0765] Server-side processing
[0766] Video Acquisition
[0767] The server acquires video data from cameras and other image acquisition devices attached to the painting robot, and transmits the video data to the server in real time, allowing the painting status to be constantly monitored.
[0768] Skin Tone Correction
[0769] Once the video data is sent to the server, the server uses the generative AI model to correct the skin color (paint color) in the video. Specifically, it uses AI image processing technology to make the color tone of the painted surface uniform.
[0770] Quality assessment
[0771] The corrected image is then further analyzed to evaluate the uniformity and texture of the paint. The server performs this evaluation using edge detection and texture analysis algorithms. The server then quantifies the evaluation results, records them in a database, and distributes them in real time.
[0772] Emotional Data Processing
[0773] The server receives video footage from cameras in the facility and generates emotion data by analyzing employees' facial expressions. The emotion engine uses facial feature points to recognize the employee's emotional state. Based on this data, the server generates feedback to help improve the efficiency of the production line and sends it to the terminal.
[0774] Terminal side processing
[0775] Video display
[0776] The terminal receives the corrected video and quality evaluation data sent from the server in real time and displays them on the screen. <video>Tag and canvas technology is used.
[0777] Viewing evaluation data
[0778] The terminal displays the evaluation data on the screen in real time, providing the user with evaluation results on the uniformity and texture of the paint.
[0779] Acquiring and sending emotion data
[0780] The device also uses a built-in camera to analyze the employee's facial expressions and generate emotional data, which is then sent to a server in real time.
[0781] User Actions
[0782] Users can operate the system using a management terminal, check the video footage and evaluation results in real time, and send feedback to the system as needed.
[0783] Specific use cases
[0784] For example, when a new painting process is introduced, a camera attached to the robot captures the painting process and sends the footage to a server. The server then uses an AI model to correct the paint color and evaluate the paint's uniformity and texture. The results are sent to the manager's device in real time. The camera also captures facial expression data from employees, and their emotional state is analyzed. Based on the analysis results, feedback to improve the efficiency of the production line is automatically provided.
[0785] Example prompt sentences to use
[0786] Here is an example of a prompt to input to a generative AI model:
[0787] Analyze employee emotions using the image below and output the results in JSON format. The facial expression data contains specific emotions including smile, anger, sadness, satisfaction, etc.
[0788] image: <base64-encoded-image-data>
[0789] Analyze the paint uniformity and texture for the following real-time video frames, quantify the evaluation results, and output them in JSON format.
[0790] Video Frame: <base64-encoded-frame-data>
[0791] As described above, the present invention provides a system that integrates the painting process and employee emotion recognition, enabling the evaluation of painting quality and the improvement of production efficiency.
[0792] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0793] Step 1:
[0794] The server acquires video data in real time from a camera attached to the painting robot. It receives the video stream sent from the camera and stores it in a buffer for processing. The input is the video data, and the output is the video frames in the buffer. Specifically, it acquires video using the camera's IP address and connection protocol (e.g., RTSP), and processes each frame using a library such as OpenCV.
[0795] Step 2:
[0796] The server inputs the captured video frame into a generative AI model to correct the paint color. The AI model analyzes each pixel of the video frame and evens out the color. The input is the video frame, and the output is a color-corrected frame. Specifically, it calls the AI model, performs inference processing, and applies the color correction algorithm.
[0797] Step 3:
[0798] The server analyzes the color-corrected video frames to evaluate the uniformity and texture of the paint. Edge detection and texture analysis algorithms are used to quantify the evaluation of the paint surface. The input is the color-corrected video frames, and the output is the numerical evaluation results. Specifically, edge detection and texture analysis are performed using OpenCV and an image analysis library, and an evaluation score is calculated.
[0799] Step 4:
[0800] The server delivers the evaluation results and corrected video frames to the administrator's terminal in real time. Real-time data transmission is performed using WebSocket or HTTP streaming. The input is the evaluation results and color-corrected video frames, and the output is the evaluation data and video displayed on the administrator's terminal. Specifically, a WebSocket server is set up and a real-time data stream is transmitted.
[0801] Step 5:
[0802] The server acquires facial expression data of employees from cameras within the facility and generates emotion data using an emotion engine. The emotion engine analyzes facial feature points and recognizes emotional states. The input is the employee's facial expression data, and the output is emotion data. Specifically, it detects faces and analyzes facial expressions to generate emotion data.
[0803] Step 6:
[0804] The server generates feedback based on the emotion data to improve the efficiency of the production line and sends it to the terminal. The input is emotion data, and the output is a feedback message. Specifically, it generates support messages and improvement suggestions based on the analysis results of the emotion data and sends them to the terminal via an HTTP request.
[0805] Step 7:
[0806] The terminal receives the corrected video and evaluation data sent from the server and displays them on the screen. The input is the evaluation data and the corrected video frame, and the output is the information displayed on the terminal's display. The specific operation is as follows: <video>Video is played using tags and canvas technology, and evaluation data is displayed using DOM manipulation.
[0807] Step 8:
[0808] The device uses a built-in camera to analyze the employee's facial expressions, generate emotional data, and send it to the server. The input is the employee's facial expression data, and the output is emotional data sent to the server. Specifically, the device acquires camera images, generates emotional data using facial expression analysis technology, and sends it to the server in real time.
[0809] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0810] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0811] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0812] [Third embodiment]
[0813] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0814] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0815] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0816] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0817] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0818] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0819] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0820] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0821] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0822] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0823] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0824] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0825] The present invention is a system for correcting the skin color of athletes in bodybuilding competitions and achieving fair muscle evaluation. The following describes the specific processing of each system element and program. The present invention is realized by the server, terminals, and users.
[0826] Server-side processing
[0827] 1. Acquiring footage
[0828] The server acquires video data from cameras and other video acquisition devices. The video is received in real-time streaming.
[0829] 2. Skin Tone Correction
[0830] The server analyzes the received video data and uses an AI model to correct the players' skin tones. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and processes them to unify their color tones.
[0831] 3. Muscle Assessment
[0832] The corrected image is further analyzed to detect muscle shadows and contours. The server then calculates a muscle evaluation score based on this. The muscle evaluation algorithm uses edge detection and depth analysis to evaluate the degree to which muscles stand out in detail.
[0833] 4. Real-time streaming
[0834] The server distributes the corrected video and muscle evaluation scores to viewers and judges in real time via the Internet or dedicated communication protocols (e.g., RTSP, WebSocket).
[0835] Terminal side processing
[0836] 1. Receiving video and evaluation data
[0837] The terminal receives the corrected video and muscle evaluation data transmitted from the server.
[0838] 2. Displaying images
[0839] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[0840] 3. Displaying evaluation data
[0841] The device displays the muscle evaluation score in real time, which is rendered as a number in a specific area using DOM manipulation.
[0842] User Actions
[0843] 1. Video selection
[0844] Users can select the video of the player they want to evaluate using the device, for example, a remote control or touch interface.
[0845] 2. Watching the video
[0846] Users can view the corrected footage and check the muscle evaluation score in real time. They can also switch to other athletes while viewing.
[0847] 3. Providing Feedback
[0848] Users can provide feedback about their viewing experience to the system, for example by entering and submitting comments and ratings using a rating form within the application.
[0849] Specific examples
[0850] Example of server operation: Video is acquired in real time, skin color is corrected using an AI model, and the corrected video is distributed along with muscle analysis results. The server repeats this process to accumulate evaluation data for each athlete.
[0851] Example of device operation: The corrected video and muscle evaluation score received from the server are displayed on the screen. When the user presses a button to switch to the video of another athlete, a different corrected video and evaluation score are displayed.
[0852] Example of user operation: Users use the application on their smartphone or tablet to select their favorite players to watch, and enjoy watching the game while referring to real-time evaluation data. After watching the game, they also submit feedback using the evaluation form in the app.
[0853] In this way, the present invention is a system that realizes a fair and transparent bodybuilding competition through video correction and muscle evaluation, thereby providing convenience and entertainment not only to athletes but also to viewers.
[0854] The processing flow will be explained below.
[0855] Server-side processing flow
[0856] Step 1:
[0857] The server acquires live video from cameras and video acquisition devices. The video signal is received in streaming format and is captured as video data in real time.
[0858] Step 2:
[0859] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[0860] Step 3:
[0861] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[0862] Step 4:
[0863] The corrected video data is then fed into another AI model to analyze muscle shading and contours. The server uses edge detection and shading analysis algorithms to assess muscle prominence.
[0864] Step 5:
[0865] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[0866] Step 6:
[0867] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[0868] Processing flow on the terminal side
[0869] Step 1:
[0870] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[0871] Step 2:
[0872] The device analyzes the received data and extracts the video data. <video>It is displayed on the screen using tags and canvases.
[0873] Step 3:
[0874] The device also receives muscle evaluation scores and displays them in a specific area on the screen. The evaluation scores are updated in real time through DOM manipulation.
[0875] User Operation Flow
[0876] Step 1:
[0877] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[0878] Step 2:
[0879] Users can view the retouched footage of their chosen athlete while viewing their muscle assessment score, which is updated in real time and displayed simultaneously.
[0880] Step 3:
[0881] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[0882] In this way, the server, terminals, and users each play their respective roles, and a system is operated that realizes a fair and transparent bodybuilding competition in real time.
[0883] Example 1
[0884] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0885] In existing bodybuilding competitions, the skin color of athletes affects the evaluation, making it difficult to provide a fair muscle evaluation. Additionally, there is a lack of systems that allow judges and viewers to provide fair evaluations in real time. Therefore, an effective method to correct athletes' skin color and provide a fair muscle evaluation is needed.
[0886] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0887] In this invention, the server includes means for acquiring video from a video acquisition device, means for correcting skin color in the acquired video, means for calculating a muscle evaluation score based on the corrected video, means for distributing the corrected video and muscle evaluation score in real time, means for displaying the received corrected video, and means for displaying the muscle evaluation score in real time. This allows for fair evaluation by correcting the skin color of the athlete, and makes it possible to provide fair muscle evaluations in real time to judges and viewers.
[0888] "Video capture device" refers to a camera or other video capture mechanism that captures video data in real time.
[0889] The "skin color correction means" is a technology that performs processing to ensure consistent skin color tones of players in the captured video.
[0890] The "means for calculating muscle evaluation scores" is a technology that analyzes the corrected video to evaluate the shading and contours of the player's muscles and convert them into a numerical score.
[0891] "Real-time distribution means" means technology that distributes the corrected footage and muscle evaluation scores to viewers and judges in real time without delay via the Internet or other communications protocols.
[0892] The "means for displaying corrected images" is a technique for displaying corrected images received from the server on the screen of the terminal.
[0893] "Means for displaying muscle evaluation scores in real time" refers to a technology that instantly displays muscle evaluation scores received from a server on the screen of a terminal.
[0894] A "generative AI model" is an artificial intelligence model that learns from large amounts of data and is designed to perform specific tasks.
[0895] The present invention is a system for correcting the skin color of athletes in bodybuilding competitions and achieving fair muscle evaluation. The present invention is realized by the server, terminals, and users.
[0896] Server-side processing
[0897] The server first acquires video data in real time from a video capture device, such as a high-resolution camera. The video is then sent to the server using the RTSP protocol, where it is received and processed by software such as FFmpeg.
[0898] The server then uses generative AI models such as TensorFlow to correct the skin tone of the captured video. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and uses techniques such as histogram matching to unify the skin tone.
[0899] The corrected video is analyzed for muscle shading and contours using OpenCV. The server performs this analysis using the Canny edge detection algorithm and depth analysis technology, and then calculates a muscle evaluation score using a proprietary algorithm based on the obtained data.
[0900] The server also uses WebSocket to deliver the corrected video and muscle evaluation scores to viewers and judges in real time. The muscle evaluation scores are overlaid on each frame, and viewers can view them in a browser or a dedicated app.
[0901] Terminal side processing
[0902] The device receives the corrected video and muscle evaluation data sent from the server using WebSocket as the reception protocol.
[0903] The device analyzes the received data and <video>The corrected image is displayed in real time using tags and Canvas technology, and the muscle evaluation score is displayed on the screen using JavaScript and DOM manipulation, allowing users to instantly check the corrected image and real-time muscle evaluation score.
[0904] User operations
[0905] The user selects the video of the desired player through the device. Using a remote control or touch interface, the desired player can be selected from a list of players on the screen. The video of the selected player is requested from the server, and the corrected video is sent to the device.
[0906] Users can check their muscle evaluation score in real time while watching the corrected video. After watching, they can provide feedback about their viewing experience using the in-app rating form. This feedback data is sent to the server and used for future analysis and improvements.
[0907] Specific examples
[0908] Example of server operation: The server acquires images in real time from a high-resolution camera and performs skin color correction using TensorFlow. Next, it analyzes muscle shading and contours using OpenCV, and delivers the corrected images and evaluation data in real time via WebSocket.
[0909] Example prompt sentence:
[0910] Capture video in real time, correct skin color using TensorFlow, analyze muscles using Canny edge detection, and deliver the results via WebSocket.
[0911] Example of device operation: The device receives the corrected video and muscle evaluation score from the server and sends them to the HTML5 <video>The tag and Canvas are used to display the player on the screen. When the user presses a button on the remote control to switch players, the new video and evaluation score are displayed.
[0912] Example prompt sentence:
[0913] Video received from the server is processed as HTML5 <video>Play with tags and use Canvas to display muscle evaluation scores. Switch players when the user selects them with the remote.
[0914] Example of user operation: The user uses the smartphone app to select their favorite player by touch operation and watch. After watching, they submit feedback using the evaluation form in the app.
[0915] Example prompt sentence:
[0916] Select a player by touching the smartphone app, and view the adjusted footage and muscle evaluation score. After viewing, submit your feedback in the evaluation form.
[0917] Thus, the present invention is a system that realizes fair and transparent bodybuilding competitions through video correction and muscle evaluation, and provides convenience and entertainment for both athletes and viewers.
[0918] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0919] Step 1:
[0920] The server acquires the video
[0921] Specifically, the server captures real-time video of the players using a high-resolution camera. This video data is sent to the server via the RTSP protocol. The input data is the raw video stream, and the output is the video data available within the server. The server receives and stores the video data using software such as FFmpeg.
[0922] Step 2:
[0923] The server corrects skin tones
[0924] After acquiring the video data, the server uses a generative AI model such as TensorFlow to perform skin tone correction. Specifically, the AI model receives the video data as input, recognizes skin-tone areas, and uses histogram matching technology to unify the color tones. The input data is the acquired video data, and the output is the corrected video data.
[0925] Step 3:
[0926] The server calculates the muscle evaluation score
[0927] The server uses OpenCV to analyze the muscle shading and contours based on the corrected video. Specifically, it uses the Canny edge detection algorithm to perform depth analysis. The input data is the corrected video data, and the output is a muscle evaluation score. The server then calculates the muscle evaluation score using a proprietary algorithm.
[0928] Step 4:
[0929] The server delivers the corrected video and muscle evaluation score in real time.
[0930] The server delivers the corrected video and muscle evaluation scores in real time via WebSocket. The muscle evaluation scores are overlaid on each frame. The input data is the corrected video and muscle evaluation scores, and the output is a stream delivered in real time. Viewers and judges can view this using a browser or a dedicated app.
[0931] Step 5:
[0932] The device receives the video and evaluation data.
[0933] The device receives the corrected video and muscle evaluation data sent from the server via WebSocket. The input data is a real-time stream from the server, and the output is the video and evaluation data available on the device. The device then sends this data to the server as HTML5 <video>Displayed using tags and Canvas technology.
[0934] Step 6:
[0935] The device displays the image
[0936] The device displays the received corrected image on the screen. <video>Video data is inserted into the tag and played back in real time. The input data is the corrected video received, and the output is the video displayed on the screen.
[0937] Step 7:
[0938] The device displays the evaluation data.
[0939] The device displays the muscle evaluation score in real time. Specifically, it analyzes the received evaluation score using JavaScript and performs DOM manipulation to draw the numerical value in a specific area. The input data is the muscle evaluation score, and the output is the evaluation score displayed on the screen.
[0940] Step 8:
[0941] The user selects a video
[0942] The user selects the video of the player they want through the terminal. Specifically, they use a remote control or touch interface to select the player they want from the player list on the screen. The input data is the user's selection information, and the output is a video request to the server.
[0943] Step 9:
[0944] The user watches the video
[0945] The user checks the muscle evaluation score in real time while watching the corrected video. Specifically, the user observes the video and evaluation score displayed on the device screen. The input data is the video and evaluation score displayed on the device, and the output is the user's visual information.
[0946] Step 10:
[0947] Users provide feedback
[0948] Users provide feedback about their viewing experience to the system. Specifically, they use the in-app rating form to enter comments and ratings and then press the submit button. The input data is the user's feedback information, and the output is feedback data sent to the server.
[0949] (Application example 1)
[0950] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0951] In conventional virtual fitting rooms, the user's skin tone is not accurately corrected, which can cause the clothes they try on to look different from how they actually look. Furthermore, there is no adequate mechanism for accurately evaluating the fit of clothes when trying them on in real time. This causes a gap between the actual fitting experience and the virtual fitting experience, reducing the reliability of the experience.
[0952] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0953] In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for overlaying a clothing model on the corrected video, means for evaluating the fit of the clothing model, and means for delivering the fit and corrected video in real time, thereby enabling accurate correction of the user's skin color, displaying the appearance of the clothing being tried on in real time, and also enabling evaluation of the fit.
[0954] "Means for inputting video" refers to a device that has the function of receiving video data from a camera or video capture device and incorporating it into the system.
[0955] The "means for correcting skin color in input video" is a device that uses an artificial intelligence model to perform processing to equalize skin tones in the acquired video and display it accurately.
[0956] The "means for overlaying a clothing model on a corrected image" is a device that performs processing to provide a realistic fitting sensation by overlaying a virtual clothing model on an image with corrected skin color.
[0957] The "means for evaluating the fit of a clothing model" is a device that performs processing to analyze and evaluate the fit of a clothing model in a corrected image.
[0958] The "means for delivering fit and corrected image in real time" refers to a device that includes a communication means and a display means for visually providing the user with the corrected image and the fit of the garment in real time.
[0959] The present invention is a system for correcting a user's skin tone in a virtual fitting room and accurately evaluating the fit of clothing. The following describes the specific processing of each element of this system and the program.
[0960] Server-side processing
[0961] The server performs the process using the following means.
[0962] 1. Acquiring footage
[0963] The server acquires video data from cameras or other video acquisition devices. The video is received in real-time streaming. The camera can be a commonly used webcam or a smartphone camera.
[0964] 2. Skin Tone Correction
[0965] The server analyzes the received video data and uses an artificial intelligence model to correct the players' skin tones. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and processes them to unify their color tones. This process uses deep learning frameworks such as TensorFlow and Keras.
[0966] 3. Overlaying the clothing model
[0967] The server overlays the virtual clothing model onto the corrected image using the OpenCV library, adjusting the position and size of the clothing to fit the user's body shape.
[0968] 4. Fit evaluation
[0969] The server evaluates the fit based on the overlaid clothing model. The fit evaluation algorithm analyzes how well the shape and size of the clothing fits the user's body type and calculates a score. This process also uses deep learning algorithms.
[0970] 5. Real-time streaming
[0971] The server delivers the corrected video and fit evaluation scores in real time using communication protocols such as WebSocket and RTSP.
[0972] Terminal side processing
[0973] The terminal receives the corrected image and fit evaluation data sent from the server and displays them to the user in the most optimal form.
[0974] 1. Receiving video and evaluation data
[0975] The terminal receives the corrected video and fit evaluation data sent from the server.
[0976] 2. Displaying images
[0977] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[0978] 3. Displaying evaluation data
[0979] The device displays the fit evaluation score in real time, which is rendered as a number in a specific area using DOM manipulation.
[0980] User Actions
[0981] Users can perform the following operations through the device:
[0982] 1. Video selection
[0983] Users can use the device to select the image of the desired garment, and then use the remote control or touchscreen interface to select the garment they want to try on.
[0984] 2. Watching the video
[0985] Users can view the corrected footage in real time and check their fit evaluation score, and can even switch to different clothing while viewing.
[0986] 3. Providing Feedback
[0987] Users can provide feedback about their viewing experience to the system, for example by entering and submitting comments and ratings using a rating form within the application.
[0988] Specific examples
[0989] Server operation example: Video is acquired in real time, skin color is corrected using an AI model, and the video is distributed with a clothing model overlaid. The server repeats this process to accumulate fit evaluation data for each garment.
[0990] Example of device operation: The corrected image and fit evaluation score received from the server are displayed on the screen. When the user presses a button to switch to the image of a different garment, a different corrected image and evaluation score are displayed.
[0991] User experience example: Users use a smartphone or tablet application to select and try on clothing items, consider purchasing them based on real-time evaluation data, and submit feedback after the try-on experience using an in-app evaluation form.
[0992] Prompt Sentence Examples
[0993] "I want to develop a virtual fitting room application that captures video in real time, corrects skin color using an AI model, and overlays selected clothing to evaluate fit. Users can capture video of themselves via smartphone or head-mounted display to see how the clothing will look on them. The program uses Python, OpenCV, and Keras."
[0994] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0995] Step 1:
[0996] The server acquires video data in real time from cameras and other video acquisition devices. Specifically, it captures the input signal from the camera and converts it into video data. This process uses the OpenCV library. The input is the video signal from the camera, and the output is the raw video data.
[0997] Step 2:
[0998] The server analyzes the acquired video data using an artificial intelligence model and performs skin color correction. Specifically, it uses a deep learning model (using TensorFlow or Keras) to recognize skin-colored areas in the video and uniformly correct their color tone. The input is RAW video data, and the output is video data with corrected skin color.
[0999] Step 3:
[1000] The server overlays a virtual clothing model on the video data with corrected skin color. Specifically, it performs a process of overlaying an image of the clothing model on the corrected video data. It performs image synthesis using the OpenCV library. The input is the video data with corrected skin color and image data of the clothing model, and the output is video data with the clothing model overlaid.
[1001] Step 4:
[1002] The server evaluates the fit of the clothing based on the overlaid video data. Using a fit evaluation algorithm (deep learning model), it analyzes how well the shape and size of the clothing in the video fits the user's body type and calculates a score. The input is the overlaid video data, and the output is a fit evaluation score.
[1003] Step 5:
[1004] The server delivers the corrected video data and fit evaluation scores in real time. Specifically, it sends the video data and evaluation scores to the user device using a distribution protocol (WebSocket or RTSP). The input is the overlaid video data and fit evaluation scores, and the output is real-time delivery to the user device.
[1005] Step 6:
[1006] The device receives the corrected video and fit evaluation data sent from the server. Specifically, it receives data from the server using WebSocket or RTSP protocols. The input is the video data and evaluation data from the server, and the output is the received data.
[1007] Step 7:
[1008] The device displays the corrected image on the screen. <video>It uses tag and canvas technology to display video in real time. The input is the received video data and the output is the video display to the user.
[1009] Step 8:
[1010] The device displays the fit evaluation score on the screen in real time. Specifically, it uses JavaScript DOM manipulation to draw the evaluation score as a numerical value in a specific area. The input is the fit evaluation data, and the output is the score displayed to the user.
[1011] Step 9:
[1012] The user selects the video of the desired garment using the device and tries it on. The streaming video is controlled using a remote control or touchscreen interface. The input is the user's selection, and the output is a video display of the selected garment.
[1013] Step 10:
[1014] After trying on the products, users provide feedback in an evaluation form within the application. Specifically, they input and submit comments and ratings through the application. The input is the user's feedback, and the output is the transmission of feedback data to the server.
[1015] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1016] This invention integrates a system that corrects the skin color of athletes in bodybuilding competitions and achieves fair muscle evaluation, with an emotion engine that recognizes the user's emotions. The following describes the specific processing of each element of the system and the program. The invention is realized by the server, terminal, and user entities.
[1017] Server-side processing
[1018] 1. Acquiring footage
[1019] The server acquires video data from cameras and other video acquisition devices, and the video data is received in real-time streaming.
[1020] 2. Skin Tone Correction
[1021] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[1022] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[1023] 3. Muscle Assessment
[1024] The corrected image is then further analyzed to detect muscle shadows and contours. The server uses edge detection and shadow analysis algorithms to evaluate muscle prominence.
[1025] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[1026] 4. Real-time streaming
[1027] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[1028] 5. Emotional Data Processing
[1029] The server receives emotional data sent from the user's device, generates feedback in real time based on that data, and sends it to the device.
[1030] Emotion engine integration
[1031] 1. Collecting Emotional Data
[1032] The device analyzes the user's facial expressions using a camera, and the emotion engine recognizes the user's emotions. The emotion engine analyzes the facial expressions using facial feature points in the video and generates emotion data.
[1033] 2. Real-time transmission
[1034] The device transmits the acquired emotion data to a server in real time using a standard protocol.
[1035] 3. Feedback Generation
[1036] The server analyzes the emotion data and generates feedback based on the user's emotion, which may provide, for example, detailed information about a specific player or a cheering message to enhance the user's viewing experience.
[1037] Terminal side processing
[1038] 1. Receiving video and evaluation data
[1039] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[1040] 2. Displaying images
[1041] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[1042] 3. Displaying evaluation data
[1043] The device displays the muscle evaluation score in real time, which is displayed in a specific area using DOM manipulation.
[1044] 4. Acquiring and sending emotion data
[1045] The device uses an emotion engine to analyze the user's facial expressions and transmits the resulting emotion data to the server.
[1046] User Actions
[1047] 1. Video selection
[1048] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[1049] 2. Watching the video
[1050] Users can view the enhanced footage of their selected athletes while viewing their muscle assessment scores, with footage and scores updated in real time.
[1051] 3. Checking emotional feedback
[1052] As users watch, they see real-time emotional feedback sent from the server, often in the form of a pop-up on their screen.
[1053] 4. Providing Feedback
[1054] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[1055] Specific examples
[1056] Example of server operation: Captures video in real time, corrects skin color using an AI model, and delivers the corrected video along with muscle analysis results. Data is also collected from the emotion engine, and feedback is generated based on that.
[1057] Example of device operation: The corrected video and muscle evaluation score received from the server are displayed on the screen, and emotional data is obtained from the user's facial expressions and sent to the server.
[1058] Example of user operation: Users use a smartphone or tablet application to select their favorite players, watch them, and enjoy watching the game while referring to real-time evaluation data and emotional feedback. After watching the game, they also submit feedback using the evaluation form in the app.
[1059] In this way, the present invention is a system that realizes fairer and more transparent bodybuilding competitions by integrating an emotion engine in addition to video correction and muscle evaluation, thereby providing convenience and entertainment not only to athletes but also to viewers.
[1060] The processing flow will be explained below.
[1061] Server-side processing flow
[1062] Step 1:
[1063] The server acquires live video from cameras and video acquisition devices. The video signal is received in streaming format and is captured as video data in real time.
[1064] Step 2:
[1065] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[1066] Step 3:
[1067] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[1068] Step 4:
[1069] The corrected video data is then fed into another AI model to analyze muscle shading and contours. The server uses edge detection and shading analysis algorithms to assess muscle prominence.
[1070] Step 5:
[1071] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[1072] Step 6:
[1073] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[1074] Step 7:
[1075] The server receives emotion data sent from the user's device, and the emotion data is transmitted to the server in real time and collected.
[1076] Step 8:
[1077] The server generates real-time feedback based on the received emotional data, and the feedback content is customized to correspond to the user's emotional state.
[1078] Step 9:
[1079] The server then sends the generated feedback to the user's device, where it is sent in real time and used to improve the viewing experience.
[1080] Processing flow on the terminal side
[1081] Step 1:
[1082] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[1083] Step 2:
[1084] The device analyzes the received data and extracts the video data. <video>It is displayed on the screen using tags and canvases.
[1085] Step 3:
[1086] The device also receives muscle evaluation scores and displays them in a specific area on the screen. The evaluation scores are updated in real time through DOM manipulation.
[1087] Step 4:
[1088] The device uses an emotion engine to analyze the user's facial expressions, processing video data from the camera and identifying the user's facial features.
[1089] Step 5:
[1090] The device recognizes the user's emotions based on the analyzed facial expression data, and the emotion engine uses an AI algorithm to determine emotions such as smile, surprise, or anger in real time.
[1091] Step 6:
[1092] The device transmits the acquired emotion data to a server in real time using a standard protocol.
[1093] Step 7:
[1094] The device receives feedback sent from the server, which is displayed on the user's screen in real time.
[1095] User Operation Flow
[1096] Step 1:
[1097] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[1098] Step 2:
[1099] Users can view the retouched footage of their chosen athlete while viewing their muscle assessment score, which is updated in real time and displayed simultaneously.
[1100] Step 3:
[1101] The user's facial expressions are analyzed in real time using the device's camera. The user does not need to perform any special operations; their natural facial expressions are automatically analyzed by the emotion engine.
[1102] Step 4:
[1103] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[1104] Step 5:
[1105] Users can view real-time feedback sent from the server, which includes detailed player information and supportive messages to enhance the viewing experience.
[1106] In this way, complex information is exchanged between the server, terminals, and users in real time, ensuring a fair and transparent bodybuilding competition.The integration of an emotion engine can make the user's viewing experience more personalized and entertaining.
[1107] Example 2
[1108] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1109] In bodybuilding competitions, differences in the skin color of athletes affect the evaluation of their muscles, making it difficult to make fair evaluations. There is also a lack of systems that can reflect viewers' emotions in real time and improve the viewing experience. The present invention aims to solve these problems and provide a system that achieves fair muscle evaluations and an improved viewing experience.
[1110] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1111] In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for evaluating muscles based on the corrected video, means for distributing the muscle evaluation and the corrected video in real time, means for collecting and analyzing emotional data, and means for generating feedback based on the analyzed emotional data. This enables fair muscle evaluation and improves the viewing experience by providing feedback that reflects the viewer's emotions.
[1112] "Means for inputting video" refers to devices and related technologies for receiving video data in real time from cameras or other video capture devices.
[1113] "Means for correcting skin color in input video" refers to devices and related technologies that use image processing technology and artificial intelligence models to adjust skin color to a standard color for received video data.
[1114] "Means for evaluating muscles based on corrected images" refers to a device and related technology that analyzes video data with corrected skin color, detects muscle shadows and contours, and calculates an evaluation score.
[1115] "Means for delivering muscle evaluation and corrected video in real time" refers to devices and related technologies for transmitting corrected video and muscle evaluation results to viewers in real time via the Internet or other communications networks.
[1116] "Means for collecting and analyzing emotional data" refers to a device and related technology that captures a user's facial expressions and movements with a camera, analyzes them, and detects their emotional state.
[1117] "Means for generating feedback based on analyzed emotional data" refers to a device and related technology for analyzing collected emotional data, creating feedback in real time based on the results, and transmitting the feedback to a viewing terminal.
[1118] This invention integrates a system that corrects the skin color of athletes in bodybuilding competitions and achieves fair muscle evaluation, as well as an emotion engine that recognizes the user's emotions. This invention is realized by the three entities: the server, the terminal, and the user.
[1119] The server acquires video data in real time using a camera or other video acquisition device. This video data is received via the RTSP protocol, for example, using FFmpeg. The received video data is then subjected to skin color correction using an artificial intelligence model (for example, TensorFlow or PyTorch). For example, a prompt such as "Perform skin color correction and apply a unified skin color model" is used as input to the AI model.
[1120] The corrected video data is then analyzed for muscle assessment. This analysis uses image processing algorithms such as Canny edge detection and Sobel filtering to detect muscle shadows and contours. A muscle assessment score is calculated based on the results of this detection. An example prompt is "Perform muscle shadow analysis and calculate an assessment score."
[1121] In addition, the server uses NGINX with the RTMP module and WebSocket to deliver the corrected video and muscle evaluation scores in real time, allowing viewers to receive unbiased video and evaluation data in real time.
[1122] The device receives and displays the corrected video and muscle evaluation score sent from the server in real time. <video>Tags and canvas technology are used to display the evaluation data on the screen using JavaScript.
[1123] The device captures the user's facial expressions and movements with a camera and uses an emotion engine (such as Face++ or Microsoft Azure Face API) to collect the user's emotion data. The collected emotion data is sent from the device to the server in real time via WebSocket. This process allows the server to constantly grasp the user's emotions and analyze them in real time.
[1124] Feedback based on emotion data is generated on the server side. For example, cheering messages and detailed information for specific players can be generated and sent to the user's device. This feedback improves the user's viewing experience. An example of a prompt is "Generate cheering messages based on emotion data."
[1125] In this way, the system of the present invention provides a fairer and more transparent evaluation by integrating an emotion engine in addition to skin color correction and muscle evaluation of the video, thereby improving the viewer experience.
[1126] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1127] Step 1:
[1128] Video Acquisition
[1129] The server receives video data in real time from the camera. The camera is installed at the bodybuilding competition venue and captures video using FFmpeg using the RTSP protocol. The input is video data as an RTSP stream, and the output is a series of video frames that can be processed on the server. This process ensures that the server always has the latest video data.
[1130] Step 2:
[1131] Skin Tone Correction
[1132] The server inputs the acquired video data into an AI model (using TensorFlow or PyTorch) to perform skin color correction. Specifically, it analyzes the color information of each pixel in the image, identifies the skin-colored areas, and performs correction. The input is the pixel data for each video frame, and the output is a video frame with the corrected skin color. As a specific example, the prompt text used is "Perform skin color correction and apply a unified skin color model." This process ensures that skin colors appear consistent even under different lighting conditions.
[1133] Step 3:
[1134] Muscle evaluation
[1135] The server analyzes the corrected video frames using an edge detection algorithm (Canny edge detection or Sobel filter) to identify muscle shading and contours. The input is the video frame with the corrected skin tone, and the output is a muscle evaluation score. Shading analysis is performed using an AI model to evaluate the prominence and shape of specific muscles. As a specific example, the prompt "Perform muscle shading analysis and calculate an evaluation score" is used. This process quantifies the definition of the athlete's muscles.
[1136] Step 4:
[1137] Real-time streaming
[1138] The server delivers the corrected video and muscle evaluation scores to viewers in real time. NGINX and the RTMP module are used to simultaneously send the video data and evaluation scores. The input is the corrected video frame and muscle evaluation score, and the output is the real-time video data and evaluation score via the distribution protocol. This processing allows viewers to receive the latest video and evaluation data in real time.
[1139] Step 5:
[1140] Collecting Emotional Data
[1141] The device captures the user's facial expressions with a camera and generates emotion data using an emotion engine (such as Face++ or Microsoft Azure Face API). The input is video data of the user's face, and the output is data indicating the user's emotional state. This process allows the user's emotions to be analyzed in real time.
[1142] Step 6:
[1143] Real-time transmission
[1144] The device transmits the generated emotion data to the server in real time via WebSocket. The input is the emotion data generated by the engine, and the output is a stream of emotion data to the server. This process allows the server to always grasp the latest emotional state of the user.
[1145] Step 7:
[1146] Feedback Generation
[1147] The server generates feedback based on the user's emotional data. The feedback includes cheering messages and detailed information about the players. The input is the analyzed emotional data, and the output is the generated feedback message. As a concrete example, the prompt "Generate cheering messages based on emotional data" is used. This process allows the user to receive information that improves the viewing experience in real time.
[1148] (Application example 2)
[1149] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1150] In conventional painting processes, the task of evaluating the uniformity and texture of the paint is manual, which can lead to unfairness and inconsistency in evaluations. Furthermore, there is a lack of mechanisms for understanding employees' working environment and emotional state in real time to optimize production efficiency. To solve these problems, a system that integrates automatic evaluation of the paint condition with employee emotional feedback is needed.
[1151] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for evaluating the quality of the paint based on the corrected video, means for distributing the paint quality evaluation and the corrected video in real time, and means for acquiring employee emotions in real time and generating feedback based on the emotion data. This automates the evaluation of paint quality, enabling fair and consistent evaluations, and furthermore, utilizing employee emotion data makes it possible to improve production line efficiency and employee satisfaction.
[1152] "Video" means visual data captured by a camera or other image capture device.
[1153] "Input means" refers to a technique for transmitting images to a server using a device such as a camera.
[1154] "Means for correcting skin tone" refers to a technology that uses an AI model to even out skin tones in video.
[1155] The "means for quality evaluation" is a technology that analyzes the uniformity and texture of the paint from the corrected image and calculates the evaluation results.
[1156] "Means for real-time distribution" refers to a technology for instantly distributing corrected video and quality evaluation results over a network.
[1157] "Means of acquiring emotions in real time" refers to technology that uses devices such as cameras to analyze employees' facial expressions and recognize their emotional state.
[1158] The "means for generating feedback based on emotional data" is a technology that uses recognized emotional data to create feedback for improving production line efficiency and employee satisfaction.
[1159] In the embodiment of the present invention, the following steps are important:
[1160] Server-side processing
[1161] Video Acquisition
[1162] The server acquires video data from cameras and other image acquisition devices attached to the painting robot, and transmits the video data to the server in real time, allowing the painting status to be constantly monitored.
[1163] Skin Tone Correction
[1164] Once the video data is sent to the server, the server uses the generative AI model to correct the skin color (paint color) in the video. Specifically, it uses AI image processing technology to make the color tone of the painted surface uniform.
[1165] Quality assessment
[1166] The corrected image is then further analyzed to evaluate the uniformity and texture of the paint. The server performs this evaluation using edge detection and texture analysis algorithms. The server then quantifies the evaluation results, records them in a database, and distributes them in real time.
[1167] Emotional Data Processing
[1168] The server receives video footage from cameras in the facility and generates emotion data by analyzing employees' facial expressions. The emotion engine uses facial feature points to recognize the employee's emotional state. Based on this data, the server generates feedback to help improve the efficiency of the production line and sends it to the terminal.
[1169] Terminal side processing
[1170] Video display
[1171] The terminal receives the corrected video and quality evaluation data sent from the server in real time and displays them on the screen. <video>Tag and canvas technology is used.
[1172] Viewing evaluation data
[1173] The terminal displays the evaluation data on the screen in real time, providing the user with evaluation results on the uniformity and texture of the paint.
[1174] Acquiring and sending emotion data
[1175] The device also uses a built-in camera to analyze the employee's facial expressions and generate emotional data, which is then sent to a server in real time.
[1176] User Actions
[1177] Users can operate the system using a management terminal, check the video footage and evaluation results in real time, and send feedback to the system as needed.
[1178] Specific use cases
[1179] For example, when a new painting process is introduced, a camera attached to the robot captures the painting process and sends the footage to a server. The server then uses an AI model to correct the paint color and evaluate the paint's uniformity and texture. The results are sent to the manager's device in real time. The camera also captures facial expression data from employees, and their emotional state is analyzed. Based on the analysis results, feedback to improve the efficiency of the production line is automatically provided.
[1180] Example prompt sentences to use
[1181] Here is an example of a prompt to input to a generative AI model:
[1182] Analyze employee emotions using the image below and output the results in JSON format. The facial expression data contains specific emotions including smile, anger, sadness, satisfaction, etc.
[1183] image: <base64-encoded-image-data>
[1184] Analyze the paint uniformity and texture for the following real-time video frames, quantify the evaluation results, and output them in JSON format.
[1185] Video Frame: <base64-encoded-frame-data>
[1186] As described above, the present invention provides a system that integrates the painting process and employee emotion recognition, enabling the evaluation of painting quality and the improvement of production efficiency.
[1187] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1188] Step 1:
[1189] The server acquires video data in real time from a camera attached to the painting robot. It receives the video stream sent from the camera and stores it in a buffer for processing. The input is the video data, and the output is the video frames in the buffer. Specifically, it acquires video using the camera's IP address and connection protocol (e.g., RTSP), and processes each frame using a library such as OpenCV.
[1190] Step 2:
[1191] The server inputs the captured video frame into a generative AI model to correct the paint color. The AI model analyzes each pixel of the video frame and evens out the color. The input is the video frame, and the output is a color-corrected frame. Specifically, it calls the AI model, performs inference processing, and applies the color correction algorithm.
[1192] Step 3:
[1193] The server analyzes the color-corrected video frames to evaluate the uniformity and texture of the paint. Edge detection and texture analysis algorithms are used to quantify the evaluation of the paint surface. The input is the color-corrected video frames, and the output is the numerical evaluation results. Specifically, edge detection and texture analysis are performed using OpenCV and an image analysis library, and an evaluation score is calculated.
[1194] Step 4:
[1195] The server delivers the evaluation results and corrected video frames to the administrator's terminal in real time. Real-time data transmission is performed using WebSocket or HTTP streaming. The input is the evaluation results and color-corrected video frames, and the output is the evaluation data and video displayed on the administrator's terminal. Specifically, a WebSocket server is set up and a real-time data stream is transmitted.
[1196] Step 5:
[1197] The server acquires facial expression data of employees from cameras within the facility and generates emotion data using an emotion engine. The emotion engine analyzes facial feature points and recognizes emotional states. The input is the employee's facial expression data, and the output is emotion data. Specifically, it detects faces and analyzes facial expressions to generate emotion data.
[1198] Step 6:
[1199] The server generates feedback based on the emotion data to improve the efficiency of the production line and sends it to the terminal. The input is emotion data, and the output is a feedback message. Specifically, it generates support messages and improvement suggestions based on the analysis results of the emotion data and sends them to the terminal via an HTTP request.
[1200] Step 7:
[1201] The terminal receives the corrected video and evaluation data sent from the server and displays them on the screen. The input is the evaluation data and the corrected video frame, and the output is the information displayed on the terminal's display. The specific operation is as follows: <video>Video is played using tags and canvas technology, and evaluation data is displayed using DOM manipulation.
[1202] Step 8:
[1203] The device uses a built-in camera to analyze the employee's facial expressions, generate emotional data, and send it to the server. The input is the employee's facial expression data, and the output is emotional data sent to the server. Specifically, the device acquires camera images, generates emotional data using facial expression analysis technology, and sends it to the server in real time.
[1204] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1205] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1206] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1207] [Fourth embodiment]
[1208] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1209] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1210] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1211] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1212] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1213] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1214] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1215] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1216] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1217] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1218] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1219] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1220] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1221] The present invention is a system for correcting the skin color of athletes in bodybuilding competitions and achieving fair muscle evaluation. The following describes the specific processing of each system element and program. The present invention is realized by the server, terminals, and users.
[1222] Server-side processing
[1223] 1. Acquiring footage
[1224] The server acquires video data from cameras and other video acquisition devices. The video is received in real-time streaming.
[1225] 2. Skin Tone Correction
[1226] The server analyzes the received video data and uses an AI model to correct the players' skin tones. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and processes them to unify their color tones.
[1227] 3. Muscle Assessment
[1228] The corrected image is further analyzed to detect muscle shadows and contours. The server then calculates a muscle evaluation score based on this. The muscle evaluation algorithm uses edge detection and depth analysis to evaluate the degree to which muscles stand out in detail.
[1229] 4. Real-time streaming
[1230] The server distributes the corrected video and muscle evaluation scores to viewers and judges in real time via the Internet or dedicated communication protocols (e.g., RTSP, WebSocket).
[1231] Terminal side processing
[1232] 1. Receiving video and evaluation data
[1233] The terminal receives the corrected video and muscle evaluation data transmitted from the server.
[1234] 2. Displaying images
[1235] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[1236] 3. Displaying evaluation data
[1237] The device displays the muscle evaluation score in real time, which is rendered as a number in a specific area using DOM manipulation.
[1238] User Actions
[1239] 1. Video selection
[1240] Users can select the video of the player they want to evaluate using the device, for example, a remote control or touch interface.
[1241] 2. Watching the video
[1242] Users can view the corrected footage and check the muscle evaluation score in real time. They can also switch to other athletes while viewing.
[1243] 3. Providing Feedback
[1244] Users can provide feedback about their viewing experience to the system, for example by entering and submitting comments and ratings using a rating form within the application.
[1245] Specific examples
[1246] Example of server operation: Video is acquired in real time, skin color is corrected using an AI model, and the corrected video is distributed along with muscle analysis results. The server repeats this process to accumulate evaluation data for each athlete.
[1247] Example of device operation: The corrected video and muscle evaluation score received from the server are displayed on the screen. When the user presses a button to switch to the video of another athlete, a different corrected video and evaluation score are displayed.
[1248] Example of user operation: Users use the application on their smartphone or tablet to select their favorite players to watch, and enjoy watching the game while referring to real-time evaluation data. After watching the game, they also submit feedback using the evaluation form in the app.
[1249] In this way, the present invention is a system that realizes a fair and transparent bodybuilding competition through video correction and muscle evaluation, thereby providing convenience and entertainment not only to athletes but also to viewers.
[1250] The processing flow will be explained below.
[1251] Server-side processing flow
[1252] Step 1:
[1253] The server acquires live video from cameras and video acquisition devices. The video signal is received in streaming format and is captured as video data in real time.
[1254] Step 2:
[1255] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[1256] Step 3:
[1257] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[1258] Step 4:
[1259] The corrected video data is then fed into another AI model to analyze muscle shading and contours. The server uses edge detection and shading analysis algorithms to assess muscle prominence.
[1260] Step 5:
[1261] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[1262] Step 6:
[1263] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[1264] Processing flow on the terminal side
[1265] Step 1:
[1266] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[1267] Step 2:
[1268] The device analyzes the received data and extracts the video data. <video>It is displayed on the screen using tags and canvases.
[1269] Step 3:
[1270] The device also receives muscle evaluation scores and displays them in a specific area on the screen. The evaluation scores are updated in real time through DOM manipulation.
[1271] User Operation Flow
[1272] Step 1:
[1273] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[1274] Step 2:
[1275] Users can view the retouched footage of their chosen athlete while viewing their muscle assessment score, which is updated in real time and displayed simultaneously.
[1276] Step 3:
[1277] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[1278] In this way, the server, terminals, and users each play their respective roles, and a system is operated that realizes a fair and transparent bodybuilding competition in real time.
[1279] Example 1
[1280] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1281] In existing bodybuilding competitions, the skin color of athletes affects the evaluation, making it difficult to provide a fair muscle evaluation. Additionally, there is a lack of systems that allow judges and viewers to provide fair evaluations in real time. Therefore, an effective method to correct athletes' skin color and provide a fair muscle evaluation is needed.
[1282] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1283] In this invention, the server includes means for acquiring video from a video acquisition device, means for correcting skin color in the acquired video, means for calculating a muscle evaluation score based on the corrected video, means for distributing the corrected video and muscle evaluation score in real time, means for displaying the received corrected video, and means for displaying the muscle evaluation score in real time. This allows for fair evaluation by correcting the skin color of the athlete, and makes it possible to provide fair muscle evaluations in real time to judges and viewers.
[1284] "Video capture device" refers to a camera or other video capture mechanism that captures video data in real time.
[1285] The "skin color correction means" is a technology that performs processing to ensure consistent skin color tones of players in the captured video.
[1286] The "means for calculating muscle evaluation scores" is a technology that analyzes the corrected video to evaluate the shading and contours of the player's muscles and convert them into a numerical score.
[1287] "Real-time distribution means" means technology that distributes the corrected footage and muscle evaluation scores to viewers and judges in real time without delay via the Internet or other communications protocols.
[1288] The "means for displaying corrected images" is a technique for displaying corrected images received from the server on the screen of the terminal.
[1289] "Means for displaying muscle evaluation scores in real time" refers to a technology that instantly displays muscle evaluation scores received from a server on the screen of a terminal.
[1290] A "generative AI model" is an artificial intelligence model that learns from large amounts of data and is designed to perform specific tasks.
[1291] The present invention is a system for correcting the skin color of athletes in bodybuilding competitions and achieving fair muscle evaluation. The present invention is realized by the server, terminals, and users.
[1292] Server-side processing
[1293] The server first acquires video data in real time from a video capture device, such as a high-resolution camera. The video is then sent to the server using the RTSP protocol, where it is received and processed by software such as FFmpeg.
[1294] The server then uses generative AI models such as TensorFlow to correct the skin tone of the captured video. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and uses techniques such as histogram matching to unify the skin tone.
[1295] The corrected video is analyzed for muscle shading and contours using OpenCV. The server performs this analysis using the Canny edge detection algorithm and depth analysis technology, and then calculates a muscle evaluation score using a proprietary algorithm based on the obtained data.
[1296] The server also uses WebSocket to deliver the corrected video and muscle evaluation scores to viewers and judges in real time. The muscle evaluation scores are overlaid on each frame, and viewers can view them in a browser or a dedicated app.
[1297] Terminal side processing
[1298] The device receives the corrected video and muscle evaluation data sent from the server using WebSocket as the reception protocol.
[1299] The device analyzes the received data and <video>The corrected image is displayed in real time using tags and Canvas technology, and the muscle evaluation score is displayed on the screen using JavaScript and DOM manipulation, allowing users to instantly check the corrected image and real-time muscle evaluation score.
[1300] User operations
[1301] The user selects the video of the desired player through the device. Using a remote control or touch interface, the desired player can be selected from a list of players on the screen. The video of the selected player is requested from the server, and the corrected video is sent to the device.
[1302] Users can check their muscle evaluation score in real time while watching the corrected video. After watching, they can provide feedback about their viewing experience using the in-app rating form. This feedback data is sent to the server and used for future analysis and improvements.
[1303] Specific examples
[1304] Example of server operation: The server acquires images in real time from a high-resolution camera and performs skin color correction using TensorFlow. Next, it analyzes muscle shading and contours using OpenCV, and delivers the corrected images and evaluation data in real time via WebSocket.
[1305] Example prompt sentence:
[1306] Capture video in real time, correct skin color using TensorFlow, analyze muscles using Canny edge detection, and deliver the results via WebSocket.
[1307] Example of device operation: The device receives the corrected video and muscle evaluation score from the server and sends them to the HTML5 <video>The tag and Canvas are used to display the player on the screen. When the user presses a button on the remote control to switch players, the new video and evaluation score are displayed.
[1308] Example prompt sentence:
[1309] Video received from the server is processed as HTML5 <video>Play with tags and use Canvas to display muscle evaluation scores. Switch players when the user selects them with the remote.
[1310] Example of user operation: The user uses the smartphone app to select their favorite player by touch operation and watch. After watching, they submit feedback using the evaluation form in the app.
[1311] Example prompt sentence:
[1312] Select a player by touching the smartphone app, and view the adjusted footage and muscle evaluation score. After viewing, submit your feedback in the evaluation form.
[1313] Thus, the present invention is a system that realizes fair and transparent bodybuilding competitions through video correction and muscle evaluation, and provides convenience and entertainment for both athletes and viewers.
[1314] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1315] Step 1:
[1316] The server acquires the video
[1317] Specifically, the server captures real-time video of the players using a high-resolution camera. This video data is sent to the server via the RTSP protocol. The input data is the raw video stream, and the output is the video data available within the server. The server receives and stores the video data using software such as FFmpeg.
[1318] Step 2:
[1319] The server corrects skin tones
[1320] After acquiring the video data, the server uses a generative AI model such as TensorFlow to perform skin tone correction. Specifically, the AI model receives the video data as input, recognizes skin-tone areas, and uses histogram matching technology to unify the color tones. The input data is the acquired video data, and the output is the corrected video data.
[1321] Step 3:
[1322] The server calculates the muscle evaluation score
[1323] The server uses OpenCV to analyze the muscle shading and contours based on the corrected video. Specifically, it uses the Canny edge detection algorithm to perform depth analysis. The input data is the corrected video data, and the output is a muscle evaluation score. The server then calculates the muscle evaluation score using a proprietary algorithm.
[1324] Step 4:
[1325] The server delivers the corrected video and muscle evaluation score in real time.
[1326] The server delivers the corrected video and muscle evaluation scores in real time via WebSocket. The muscle evaluation scores are overlaid on each frame. The input data is the corrected video and muscle evaluation scores, and the output is a stream delivered in real time. Viewers and judges can view this using a browser or a dedicated app.
[1327] Step 5:
[1328] The device receives the video and evaluation data.
[1329] The device receives the corrected video and muscle evaluation data sent from the server via WebSocket. The input data is a real-time stream from the server, and the output is the video and evaluation data available on the device. The device then sends this data to the server as HTML5 <video>Displayed using tags and Canvas technology.
[1330] Step 6:
[1331] The device displays the image
[1332] The device displays the received corrected image on the screen. <video>Video data is inserted into the tag and played back in real time. The input data is the corrected video received, and the output is the video displayed on the screen.
[1333] Step 7:
[1334] The device displays the evaluation data.
[1335] The device displays the muscle evaluation score in real time. Specifically, it analyzes the received evaluation score using JavaScript and performs DOM manipulation to draw the numerical value in a specific area. The input data is the muscle evaluation score, and the output is the evaluation score displayed on the screen.
[1336] Step 8:
[1337] The user selects a video
[1338] The user selects the video of the player they want through the terminal. Specifically, they use a remote control or touch interface to select the player they want from the player list on the screen. The input data is the user's selection information, and the output is a video request to the server.
[1339] Step 9:
[1340] The user watches the video
[1341] The user checks the muscle evaluation score in real time while watching the corrected video. Specifically, the user observes the video and evaluation score displayed on the device screen. The input data is the video and evaluation score displayed on the device, and the output is the user's visual information.
[1342] Step 10:
[1343] Users provide feedback
[1344] Users provide feedback about their viewing experience to the system. Specifically, they use the in-app rating form to enter comments and ratings and then press the submit button. The input data is the user's feedback information, and the output is feedback data sent to the server.
[1345] (Application example 1)
[1346] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1347] In conventional virtual fitting rooms, the user's skin tone is not accurately corrected, which can cause the clothes they try on to look different from how they actually look. Furthermore, there is no adequate mechanism for accurately evaluating the fit of clothes when trying them on in real time. This causes a gap between the actual fitting experience and the virtual fitting experience, reducing the reliability of the experience.
[1348] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1349] In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for overlaying a clothing model on the corrected video, means for evaluating the fit of the clothing model, and means for delivering the fit and corrected video in real time, thereby enabling accurate correction of the user's skin color, displaying the appearance of the clothing being tried on in real time, and also enabling evaluation of the fit.
[1350] "Means for inputting video" refers to a device that has the function of receiving video data from a camera or video capture device and incorporating it into the system.
[1351] The "means for correcting skin color in input video" is a device that uses an artificial intelligence model to perform processing to equalize skin tones in the acquired video and display it accurately.
[1352] The "means for overlaying a clothing model on a corrected image" is a device that performs processing to provide a realistic fitting sensation by overlaying a virtual clothing model on an image with corrected skin color.
[1353] The "means for evaluating the fit of a clothing model" is a device that performs processing to analyze and evaluate the fit of a clothing model in a corrected image.
[1354] The "means for delivering fit and corrected image in real time" refers to a device that includes a communication means and a display means for visually providing the user with the corrected image and the fit of the garment in real time.
[1355] The present invention is a system for correcting a user's skin tone in a virtual fitting room and accurately evaluating the fit of clothing. The following describes the specific processing of each element of this system and the program.
[1356] Server-side processing
[1357] The server performs the process using the following means.
[1358] 1. Acquiring footage
[1359] The server acquires video data from cameras or other video acquisition devices. The video is received in real-time streaming. The camera can be a commonly used webcam or a smartphone camera.
[1360] 2. Skin Tone Correction
[1361] The server analyzes the received video data and uses an artificial intelligence model to correct the players' skin tones. Specifically, the skin tone correction algorithm recognizes skin-colored areas in the video and processes them to unify their color tones. This process uses deep learning frameworks such as TensorFlow and Keras.
[1362] 3. Overlaying the clothing model
[1363] The server overlays the virtual clothing model onto the corrected image using the OpenCV library, adjusting the position and size of the clothing to fit the user's body shape.
[1364] 4. Fit evaluation
[1365] The server evaluates the fit based on the overlaid clothing model. The fit evaluation algorithm analyzes how well the shape and size of the clothing fits the user's body type and calculates a score. This process also uses deep learning algorithms.
[1366] 5. Real-time streaming
[1367] The server delivers the corrected video and fit evaluation scores in real time using communication protocols such as WebSocket and RTSP.
[1368] Terminal side processing
[1369] The terminal receives the corrected image and fit evaluation data sent from the server and displays them to the user in the most optimal form.
[1370] 1. Receiving video and evaluation data
[1371] The terminal receives the corrected video and fit evaluation data sent from the server.
[1372] 2. Displaying images
[1373] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[1374] 3. Displaying evaluation data
[1375] The device displays the fit evaluation score in real time, which is rendered as a number in a specific area using DOM manipulation.
[1376] User Actions
[1377] Users can perform the following operations through the device:
[1378] 1. Video selection
[1379] Users can use the device to select the image of the desired garment, and then use the remote control or touchscreen interface to select the garment they want to try on.
[1380] 2. Watching the video
[1381] Users can view the corrected footage in real time and check their fit evaluation score, and can even switch to different clothing while viewing.
[1382] 3. Providing Feedback
[1383] Users can provide feedback about their viewing experience to the system, for example by entering and submitting comments and ratings using a rating form within the application.
[1384] Specific examples
[1385] Server operation example: Video is acquired in real time, skin color is corrected using an AI model, and the video is distributed with a clothing model overlaid. The server repeats this process to accumulate fit evaluation data for each garment.
[1386] Example of device operation: The corrected image and fit evaluation score received from the server are displayed on the screen. When the user presses a button to switch to the image of a different garment, a different corrected image and evaluation score are displayed.
[1387] User experience example: Users use a smartphone or tablet application to select and try on clothing items, consider purchasing them based on real-time evaluation data, and submit feedback after the try-on experience using an in-app evaluation form.
[1388] Prompt Sentence Examples
[1389] "I want to develop a virtual fitting room application that captures video in real time, corrects skin color using an AI model, and overlays selected clothing to evaluate fit. Users can capture video of themselves via smartphone or head-mounted display to see how the clothing will look on them. The program uses Python, OpenCV, and Keras."
[1390] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1391] Step 1:
[1392] The server acquires video data in real time from cameras and other video acquisition devices. Specifically, it captures the input signal from the camera and converts it into video data. This process uses the OpenCV library. The input is the video signal from the camera, and the output is the raw video data.
[1393] Step 2:
[1394] The server analyzes the acquired video data using an artificial intelligence model and performs skin color correction. Specifically, it uses a deep learning model (using TensorFlow or Keras) to recognize skin-colored areas in the video and uniformly correct their color tone. The input is RAW video data, and the output is video data with corrected skin color.
[1395] Step 3:
[1396] The server overlays a virtual clothing model on the video data with corrected skin color. Specifically, it performs a process of overlaying an image of the clothing model on the corrected video data. It performs image synthesis using the OpenCV library. The input is the video data with corrected skin color and image data of the clothing model, and the output is video data with the clothing model overlaid.
[1397] Step 4:
[1398] The server evaluates the fit of the clothing based on the overlaid video data. Using a fit evaluation algorithm (deep learning model), it analyzes how well the shape and size of the clothing in the video fits the user's body type and calculates a score. The input is the overlaid video data, and the output is a fit evaluation score.
[1399] Step 5:
[1400] The server delivers the corrected video data and fit evaluation scores in real time. Specifically, it sends the video data and evaluation scores to the user device using a distribution protocol (WebSocket or RTSP). The input is the overlaid video data and fit evaluation scores, and the output is real-time delivery to the user device.
[1401] Step 6:
[1402] The device receives the corrected video and fit evaluation data sent from the server. Specifically, it receives data from the server using WebSocket or RTSP protocols. The input is the video data and evaluation data from the server, and the output is the received data.
[1403] Step 7:
[1404] The device displays the corrected image on the screen. <video>It uses tag and canvas technology to display video in real time. The input is the received video data and the output is the video display to the user.
[1405] Step 8:
[1406] The device displays the fit evaluation score on the screen in real time. Specifically, it uses JavaScript DOM manipulation to draw the evaluation score as a numerical value in a specific area. The input is the fit evaluation data, and the output is the score displayed to the user.
[1407] Step 9:
[1408] The user selects the video of the desired garment using the device and tries it on. The streaming video is controlled using a remote control or touchscreen interface. The input is the user's selection, and the output is a video display of the selected garment.
[1409] Step 10:
[1410] After trying on the products, users provide feedback in an evaluation form within the application. Specifically, they input and submit comments and ratings through the application. The input is the user's feedback, and the output is the transmission of feedback data to the server.
[1411] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1412] This invention integrates a system that corrects the skin color of athletes in bodybuilding competitions and achieves fair muscle evaluation, with an emotion engine that recognizes the user's emotions. The following describes the specific processing of each element of the system and the program. The invention is realized by the server, terminal, and user entities.
[1413] Server-side processing
[1414] 1. Acquiring footage
[1415] The server acquires video data from cameras and other video acquisition devices, and the video data is received in real-time streaming.
[1416] 2. Skin Tone Correction
[1417] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[1418] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[1419] 3. Muscle Assessment
[1420] The corrected image is then further analyzed to detect muscle shadows and contours. The server uses edge detection and shadow analysis algorithms to evaluate muscle prominence.
[1421] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[1422] 4. Real-time streaming
[1423] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[1424] 5. Emotional Data Processing
[1425] The server receives emotional data sent from the user's device, generates feedback in real time based on that data, and sends it to the device.
[1426] Emotion engine integration
[1427] 1. Collecting Emotional Data
[1428] The device analyzes the user's facial expressions using a camera, and the emotion engine recognizes the user's emotions. The emotion engine analyzes the facial expressions using facial feature points in the video and generates emotion data.
[1429] 2. Real-time transmission
[1430] The device transmits the acquired emotion data to a server in real time using a standard protocol.
[1431] 3. Feedback Generation
[1432] The server analyzes the emotion data and generates feedback based on the user's emotion, which may provide, for example, detailed information about a specific player or a cheering message to enhance the user's viewing experience.
[1433] Terminal side processing
[1434] 1. Receiving video and evaluation data
[1435] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[1436] 2. Displaying images
[1437] The device displays the corrected image on the screen. <video>Tag and canvas technology is used.
[1438] 3. Displaying evaluation data
[1439] The device displays the muscle evaluation score in real time, which is displayed in a specific area using DOM manipulation.
[1440] 4. Acquiring and sending emotion data
[1441] The device uses an emotion engine to analyze the user's facial expressions and transmits the resulting emotion data to the server.
[1442] User Actions
[1443] 1. Video selection
[1444] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[1445] 2. Watching the video
[1446] Users can view the enhanced footage of their selected athletes while viewing their muscle assessment scores, with footage and scores updated in real time.
[1447] 3. Checking emotional feedback
[1448] As users watch, they see real-time emotional feedback sent from the server, often in the form of a pop-up on their screen.
[1449] 4. Providing Feedback
[1450] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[1451] Specific examples
[1452] Example of server operation: Captures video in real time, corrects skin color using an AI model, and delivers the corrected video along with muscle analysis results. Data is also collected from the emotion engine, and feedback is generated based on that.
[1453] Example of device operation: The corrected video and muscle evaluation score received from the server are displayed on the screen, and emotional data is obtained from the user's facial expressions and sent to the server.
[1454] Example of user operation: Users use a smartphone or tablet application to select their favorite players, watch them, and enjoy watching the game while referring to real-time evaluation data and emotional feedback. After watching the game, they also submit feedback using the evaluation form in the app.
[1455] In this way, the present invention is a system that realizes fairer and more transparent bodybuilding competitions by integrating an emotion engine in addition to video correction and muscle evaluation, thereby providing convenience and entertainment not only to athletes but also to viewers.
[1456] The processing flow will be explained below.
[1457] Server-side processing flow
[1458] Step 1:
[1459] The server acquires live video from cameras and video acquisition devices. The video signal is received in streaming format and is captured as video data in real time.
[1460] Step 2:
[1461] The server inputs the received video data into the AI model and identifies skin-colored areas within the video. Specifically, it uses image processing technology to extract skin-colored areas.
[1462] Step 3:
[1463] The server uses an AI algorithm to perform color correction on the identified skin tone areas, standardizing skin tones and ensuring a consistent look across the entire image.
[1464] Step 4:
[1465] The corrected video data is then fed into another AI model to analyze muscle shading and contours. The server uses edge detection and shading analysis algorithms to assess muscle prominence.
[1466] Step 5:
[1467] Based on the analysis results, the server calculates a muscle evaluation score for each athlete. This score is converted into a numerical value and then sent to the next processing step along with the video.
[1468] Step 6:
[1469] The server delivers the corrected video and calculated muscle evaluation scores in real time using WebSocket or RTSP as the communication protocol.
[1470] Step 7:
[1471] The server receives emotion data sent from the user's device, and the emotion data is transmitted to the server in real time and collected.
[1472] Step 8:
[1473] The server generates real-time feedback based on the received emotional data, and the feedback content is customized to correspond to the user's emotional state.
[1474] Step 9:
[1475] The server then sends the generated feedback to the user's device, where it is sent in real time and used to improve the viewing experience.
[1476] Processing flow on the terminal side
[1477] Step 1:
[1478] The device receives the corrected video and muscle evaluation data sent from the server in real time. The device opens a WebSocket connection to secure the data stream.
[1479] Step 2:
[1480] The device analyzes the received data and extracts the video data. <video>It is displayed on the screen using tags and canvases.
[1481] Step 3:
[1482] The device also receives muscle evaluation scores and displays them in a specific area on the screen. The evaluation scores are updated in real time through DOM manipulation.
[1483] Step 4:
[1484] The device uses an emotion engine to analyze the user's facial expressions, processing video data from the camera and identifying the user's facial features.
[1485] Step 5:
[1486] The device recognizes the user's emotions based on the analyzed facial expression data, and the emotion engine uses an AI algorithm to determine emotions such as smile, surprise, or anger in real time.
[1487] Step 6:
[1488] The device transmits the acquired emotion data to a server in real time using a standard protocol.
[1489] Step 7:
[1490] The device receives feedback sent from the server, which is displayed on the user's screen in real time.
[1491] User Operation Flow
[1492] Step 1:
[1493] Users can use a remote control or touch interface to select the player image they want to display by selecting the desired player from the menu on the UI and pressing the confirm button.
[1494] Step 2:
[1495] Users can view the retouched footage of their chosen athlete while viewing their muscle assessment score, which is updated in real time and displayed simultaneously.
[1496] Step 3:
[1497] The user's facial expressions are analyzed in real time using the device's camera. The user does not need to perform any special operations; their natural facial expressions are automatically analyzed by the emotion engine.
[1498] Step 4:
[1499] Users can submit feedback about their viewing experience through a form within the app. Feedback is entered in text format and sent to the server by pressing the submit button.
[1500] Step 5:
[1501] Users can view real-time feedback sent from the server, which includes detailed player information and supportive messages to enhance the viewing experience.
[1502] In this way, complex information is exchanged between the server, terminals, and users in real time, ensuring a fair and transparent bodybuilding competition.The integration of an emotion engine can make the user's viewing experience more personalized and entertaining.
[1503] Example 2
[1504] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1505] In bodybuilding competitions, differences in the skin color of athletes affect the evaluation of their muscles, making it difficult to make fair evaluations. There is also a lack of systems that can reflect viewers' emotions in real time and improve the viewing experience. The present invention aims to solve these problems and provide a system that achieves fair muscle evaluations and an improved viewing experience.
[1506] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1507] In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for evaluating muscles based on the corrected video, means for distributing the muscle evaluation and the corrected video in real time, means for collecting and analyzing emotional data, and means for generating feedback based on the analyzed emotional data. This enables fair muscle evaluation and improves the viewing experience by providing feedback that reflects the viewer's emotions.
[1508] "Means for inputting video" refers to devices and related technologies for receiving video data in real time from cameras or other video capture devices.
[1509] "Means for correcting skin color in input video" refers to devices and related technologies that use image processing technology and artificial intelligence models to adjust skin color to a standard color for received video data.
[1510] "Means for evaluating muscles based on corrected images" refers to a device and related technology that analyzes video data with corrected skin color, detects muscle shadows and contours, and calculates an evaluation score.
[1511] "Means for delivering muscle evaluation and corrected video in real time" refers to devices and related technologies for transmitting corrected video and muscle evaluation results to viewers in real time via the Internet or other communications networks.
[1512] "Means for collecting and analyzing emotional data" refers to a device and related technology that captures a user's facial expressions and movements with a camera, analyzes them, and detects their emotional state.
[1513] "Means for generating feedback based on analyzed emotional data" refers to a device and related technology for analyzing collected emotional data, creating feedback in real time based on the results, and transmitting the feedback to a viewing terminal.
[1514] This invention integrates a system that corrects the skin color of athletes in bodybuilding competitions and achieves fair muscle evaluation, as well as an emotion engine that recognizes the user's emotions. This invention is realized by the three entities: the server, the terminal, and the user.
[1515] The server acquires video data in real time using a camera or other video acquisition device. This video data is received via the RTSP protocol, for example, using FFmpeg. The received video data is then subjected to skin color correction using an artificial intelligence model (for example, TensorFlow or PyTorch). For example, a prompt such as "Perform skin color correction and apply a unified skin color model" is used as input to the AI model.
[1516] The corrected video data is then analyzed for muscle assessment. This analysis uses image processing algorithms such as Canny edge detection and Sobel filtering to detect muscle shadows and contours. A muscle assessment score is calculated based on the results of this detection. An example prompt is "Perform muscle shadow analysis and calculate an assessment score."
[1517] In addition, the server uses NGINX with the RTMP module and WebSocket to deliver the corrected video and muscle evaluation scores in real time, allowing viewers to receive unbiased video and evaluation data in real time.
[1518] The device receives and displays the corrected video and muscle evaluation score sent from the server in real time. <video>Tags and canvas technology are used to display the evaluation data on the screen using JavaScript.
[1519] The device captures the user's facial expressions and movements with a camera and uses an emotion engine (such as Face++ or Microsoft Azure Face API) to collect the user's emotion data. The collected emotion data is sent from the device to the server in real time via WebSocket. This process allows the server to constantly grasp the user's emotions and analyze them in real time.
[1520] Feedback based on emotion data is generated on the server side. For example, cheering messages and detailed information for specific players can be generated and sent to the user's device. This feedback improves the user's viewing experience. An example of a prompt is "Generate cheering messages based on emotion data."
[1521] In this way, the system of the present invention provides a fairer and more transparent evaluation by integrating an emotion engine in addition to skin color correction and muscle evaluation of the video, thereby improving the viewer experience.
[1522] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1523] Step 1:
[1524] Video Acquisition
[1525] The server receives video data in real time from the camera. The camera is installed at the bodybuilding competition venue and captures video using FFmpeg using the RTSP protocol. The input is video data as an RTSP stream, and the output is a series of video frames that can be processed on the server. This process ensures that the server always has the latest video data.
[1526] Step 2:
[1527] Skin Tone Correction
[1528] The server inputs the acquired video data into an AI model (using TensorFlow or PyTorch) to perform skin color correction. Specifically, it analyzes the color information of each pixel in the image, identifies the skin-colored areas, and performs correction. The input is the pixel data for each video frame, and the output is a video frame with the corrected skin color. As a specific example, the prompt text used is "Perform skin color correction and apply a unified skin color model." This process ensures that skin colors appear consistent even under different lighting conditions.
[1529] Step 3:
[1530] Muscle evaluation
[1531] The server analyzes the corrected video frames using an edge detection algorithm (Canny edge detection or Sobel filter) to identify muscle shading and contours. The input is the video frame with the corrected skin tone, and the output is a muscle evaluation score. Shading analysis is performed using an AI model to evaluate the prominence and shape of specific muscles. As a specific example, the prompt "Perform muscle shading analysis and calculate an evaluation score" is used. This process quantifies the definition of the athlete's muscles.
[1532] Step 4:
[1533] Real-time streaming
[1534] The server delivers the corrected video and muscle evaluation scores to viewers in real time. NGINX and the RTMP module are used to simultaneously send the video data and evaluation scores. The input is the corrected video frame and muscle evaluation score, and the output is the real-time video data and evaluation score via the distribution protocol. This processing allows viewers to receive the latest video and evaluation data in real time.
[1535] Step 5:
[1536] Collecting Emotional Data
[1537] The device captures the user's facial expressions with a camera and generates emotion data using an emotion engine (such as Face++ or Microsoft Azure Face API). The input is video data of the user's face, and the output is data indicating the user's emotional state. This process allows the user's emotions to be analyzed in real time.
[1538] Step 6:
[1539] Real-time transmission
[1540] The device transmits the generated emotion data to the server in real time via WebSocket. The input is the emotion data generated by the engine, and the output is a stream of emotion data to the server. This process allows the server to always grasp the latest emotional state of the user.
[1541] Step 7:
[1542] Feedback Generation
[1543] The server generates feedback based on the user's emotional data. The feedback includes cheering messages and detailed information about the players. The input is the analyzed emotional data, and the output is the generated feedback message. As a concrete example, the prompt "Generate cheering messages based on emotional data" is used. This process allows the user to receive information that improves the viewing experience in real time.
[1544] (Application example 2)
[1545] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1546] In conventional painting processes, the task of evaluating the uniformity and texture of the paint is manual, which can lead to unfairness and inconsistency in evaluations. Furthermore, there is a lack of mechanisms for understanding employees' working environment and emotional state in real time to optimize production efficiency. To solve these problems, a system that integrates automatic evaluation of the paint condition with employee emotional feedback is needed.
[1547] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting video, means for correcting skin color in the input video, means for evaluating the quality of the paint based on the corrected video, means for distributing the paint quality evaluation and the corrected video in real time, and means for acquiring employee emotions in real time and generating feedback based on the emotion data. This automates the evaluation of paint quality, enabling fair and consistent evaluations, and furthermore, utilizing employee emotion data makes it possible to improve production line efficiency and employee satisfaction.
[1548] "Video" means visual data captured by a camera or other image capture device.
[1549] "Input means" refers to a technique for transmitting images to a server using a device such as a camera.
[1550] "Means for correcting skin tone" refers to a technology that uses an AI model to even out skin tones in video.
[1551] The "means for quality evaluation" is a technology that analyzes the uniformity and texture of the paint from the corrected image and calculates the evaluation results.
[1552] "Means for real-time distribution" refers to a technology for instantly distributing corrected video and quality evaluation results over a network.
[1553] "Means of acquiring emotions in real time" refers to technology that uses devices such as cameras to analyze employees' facial expressions and recognize their emotional state.
[1554] The "means for generating feedback based on emotional data" is a technology that uses recognized emotional data to create feedback for improving production line efficiency and employee satisfaction.
[1555] In the embodiment of the present invention, the following steps are important:
[1556] Server-side processing
[1557] Video Acquisition
[1558] The server acquires video data from cameras and other image acquisition devices attached to the painting robot, and transmits the video data to the server in real time, allowing the painting status to be constantly monitored.
[1559] Skin Tone Correction
[1560] Once the video data is sent to the server, the server uses the generative AI model to correct the skin color (paint color) in the video. Specifically, it uses AI image processing technology to make the color tone of the painted surface uniform.
[1561] Quality assessment
[1562] The corrected image is then further analyzed to evaluate the uniformity and texture of the paint. The server performs this evaluation using edge detection and texture analysis algorithms. The server then quantifies the evaluation results, records them in a database, and distributes them in real time.
[1563] Emotional Data Processing
[1564] The server receives video footage from cameras in the facility and generates emotion data by analyzing employees' facial expressions. The emotion engine uses facial feature points to recognize the employee's emotional state. Based on this data, the server generates feedback to help improve the efficiency of the production line and sends it to the terminal.
[1565] Terminal side processing
[1566] Video display
[1567] The terminal receives the corrected video and quality evaluation data sent from the server in real time and displays them on the screen. <video>Tag and canvas technology is used.
[1568] Viewing evaluation data
[1569] The terminal displays the evaluation data on the screen in real time, providing the user with evaluation results on the uniformity and texture of the paint.
[1570] Acquiring and sending emotion data
[1571] The device also uses a built-in camera to analyze the employee's facial expressions and generate emotional data, which is then sent to a server in real time.
[1572] User Actions
[1573] Users can operate the system using a management terminal, check the video footage and evaluation results in real time, and send feedback to the system as needed.
[1574] Specific use cases
[1575] For example, when a new painting process is introduced, a camera attached to the robot captures the painting process and sends the footage to a server. The server then uses an AI model to correct the paint color and evaluate the paint's uniformity and texture. The results are sent to the manager's device in real time. The camera also captures facial expression data from employees, and their emotional state is analyzed. Based on the analysis results, feedback to improve the efficiency of the production line is automatically provided.
[1576] Example prompt sentences to use
[1577] Here is an example of a prompt to input to a generative AI model:
[1578] Analyze employee emotions using the image below and output the results in JSON format. The facial expression data contains specific emotions including smile, anger, sadness, satisfaction, etc.
[1579] image: <base64-encoded-image-data>
[1580] Analyze the paint uniformity and texture for the following real-time video frames, quantify the evaluation results, and output them in JSON format.
[1581] Video Frame: <base64-encoded-frame-data>
[1582] As described above, the present invention provides a system that integrates the painting process and employee emotion recognition, enabling the evaluation of painting quality and the improvement of production efficiency.
[1583] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1584] Step 1:
[1585] The server acquires video data in real time from a camera attached to the painting robot. It receives the video stream sent from the camera and stores it in a buffer for processing. The input is the video data, and the output is the video frames in the buffer. Specifically, it acquires video using the camera's IP address and connection protocol (e.g., RTSP), and processes each frame using a library such as OpenCV.
[1586] Step 2:
[1587] The server inputs the captured video frame into a generative AI model to correct the paint color. The AI model analyzes each pixel of the video frame and evens out the color. The input is the video frame, and the output is a color-corrected frame. Specifically, it calls the AI model, performs inference processing, and applies the color correction algorithm.
[1588] Step 3:
[1589] The server analyzes the color-corrected video frames to evaluate the uniformity and texture of the paint. Edge detection and texture analysis algorithms are used to quantify the evaluation of the paint surface. The input is the color-corrected video frames, and the output is the numerical evaluation results. Specifically, edge detection and texture analysis are performed using OpenCV and an image analysis library, and an evaluation score is calculated.
[1590] Step 4:
[1591] The server delivers the evaluation results and corrected video frames to the administrator's terminal in real time. Real-time data transmission is performed using WebSocket or HTTP streaming. The input is the evaluation results and color-corrected video frames, and the output is the evaluation data and video displayed on the administrator's terminal. Specifically, a WebSocket server is set up and a real-time data stream is transmitted.
[1592] Step 5:
[1593] The server acquires facial expression data of employees from cameras within the facility and generates emotion data using an emotion engine. The emotion engine analyzes facial feature points and recognizes emotional states. The input is the employee's facial expression data, and the output is emotion data. Specifically, it detects faces and analyzes facial expressions to generate emotion data.
[1594] Step 6:
[1595] The server generates feedback based on the emotion data to improve the efficiency of the production line and sends it to the terminal. The input is emotion data, and the output is a feedback message. Specifically, it generates support messages and improvement suggestions based on the analysis results of the emotion data and sends them to the terminal via an HTTP request.
[1596] Step 7:
[1597] The terminal receives the corrected video and evaluation data sent from the server and displays them on the screen. The input is the evaluation data and the corrected video frame, and the output is the information displayed on the terminal's display. The specific operation is as follows: <video>Video is played using tags and canvas technology, and evaluation data is displayed using DOM manipulation.
[1598] Step 8:
[1599] The device uses a built-in camera to analyze the employee's facial expressions, generate emotional data, and send it to the server. The input is the employee's facial expression data, and the output is emotional data sent to the server. Specifically, the device acquires camera images, generates emotional data using facial expression analysis technology, and sends it to the server in real time.
[1600] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1601] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1602] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1603] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1604] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1605] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1606] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1607] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1608] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1609] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1610] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1611] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1612] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1613] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1614] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1615] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1616] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1617] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1618] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1619] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1620] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1621] The following is further disclosed regarding the above embodiment.
[1622] (Claim 1)
[1623] A means for inputting video;
[1624] means for correcting skin color in the input video;
[1625] A means for performing muscle evaluation based on the corrected image;
[1626] A means for delivering muscle evaluation and corrected footage in real time;
[1627] A system including:
[1628] (Claim 2)
[1629] 10. The system of claim 1, wherein the skin color correction means uses an artificial intelligence model to process the skin tones uniformly.
[1630] (Claim 3)
[1631] 10. The system of claim 1, wherein the muscle evaluation means performs processing to analyze muscle shadows and contours from the corrected image.
[1632] "Example 1"
[1633] (Claim 1)
[1634] means for acquiring an image from an image acquisition device;
[1635] means for correcting skin tones in the captured video;
[1636] A means for calculating a muscle evaluation score based on the corrected video;
[1637] a means for delivering the corrected video and muscle assessment scores in real time;
[1638] means for displaying the received corrected image;
[1639] a means for displaying muscle assessment scores in real time;
[1640] A system including:
[1641] (Claim 2)
[1642] 2. The system of claim 1, wherein the skin color correction means uses a generative AI model to perform processing to unify skin tones.
[1643] (Claim 3)
[1644] 2. The system according to claim 1, wherein the muscle evaluation means performs processing to analyze muscle shadows and contours from the corrected image using edge detection and depth analysis.
[1645] "Application Example 1"
[1646] (Claim 1)
[1647] A means for inputting video;
[1648] means for correcting skin color in the input video;
[1649] means for overlaying a garment model on the corrected image;
[1650] a means for assessing the fit of a clothing model;
[1651] a means for delivering the fitted and corrected footage in real time;
[1652] A system including:
[1653] (Claim 2)
[1654] 10. The system of claim 1, wherein the skin color correction means uses an artificial intelligence model to process the skin tones uniformly.
[1655] (Claim 3)
[1656] 2. The system according to claim 1, wherein the clothing model overlay means performs processing to overlay the clothing model on the corrected video in real time.
[1657] "Example 2: Combining Emotion Engines"
[1658] (Claim 1)
[1659] A means for inputting video;
[1660] means for correcting skin color in the input video;
[1661] A means for performing muscle evaluation based on the corrected image;
[1662] A means for delivering muscle evaluation and corrected footage in real time;
[1663] a means for collecting and analyzing emotion data;
[1664] means for generating feedback based on the analyzed emotion data;
[1665] A system including:
[1666] (Claim 2)
[1667] 10. The system of claim 1, wherein the skin color correction means uses an artificial intelligence model to process the skin tones uniformly.
[1668] (Claim 3)
[1669] 10. The system of claim 1, wherein the muscle evaluation means performs processing to analyze muscle shadows and contours from the corrected image.
[1670] (Claim 4)
[1671] 2. The system according to claim 1, wherein the means for collecting emotion data performs processing to analyze the user's facial expressions with a camera and generate emotion data using an emotion engine.
[1672] (Claim 5)
[1673] 2. The system according to claim 1, wherein the feedback generating means performs processing to generate feedback in real time based on the analyzed emotion data and transmit the feedback to the terminal.
[1674] "Application example 2 when combining emotion engines"
[1675] (Claim 1)
[1676] A means for inputting video;
[1677] means for correcting skin color in the input video;
[1678] A means for evaluating the quality of the coating based on the corrected image;
[1679] A means for delivering paint quality evaluation and corrected images in real time;
[1680] A means of capturing employee sentiment in real time and generating feedback based on that sentiment data;
[1681] A system including:
[1682] (Claim 2)
[1683] 10. The system of claim 1, wherein the skin color correction means uses an artificial intelligence model to process the skin tones uniformly.
[1684] (Claim 3)
[1685] 2. The system of claim 1, wherein the paint quality evaluation means performs processing to analyze the uniformity and texture of the paint from the corrected image. [Explanation of symbols]
[1686] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / url:> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / url:> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / url:> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video> < / video>
Claims
1. A means for inputting video; means for correcting skin color in the input video; A means for performing muscle evaluation based on the corrected image; A means for delivering muscle evaluation and corrected footage in real time; A system including:
2. 2. The system of claim 1, wherein the skin color correction means uses an artificial intelligence model to process the skin tones uniformly.
3. 2. The system of claim 1, wherein the muscle evaluation means performs processing to analyze muscle shadows and contours from the corrected image.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A