System
A system that analyzes and overlays technical and compositional scores in real-time on sports video feeds addresses spectator confusion, improving understanding and enjoyment.
Patent Information
- Application Number
- JP2024116437
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Spectators often struggle to understand scoring criteria in sports events like figure skating, ski jumping, and snowboarding, leading to confusion and reduced enjoyment due to unclear technical and compositional scores and deductions.
A system that receives a video feed, analyzes each frame to generate technical scores, composition scores, and deductions, and overlays these results in real-time on the video for immediate visual confirmation.
Enhances transparency and enjoyment of sports viewing by allowing spectators to intuitively understand scoring criteria in real-time.
Smart Images

Figure 2026014963000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When watching sports, confusing scoring criteria is often a problem. Particularly in events such as figure skating, ski jumping, and snowboarding, it is difficult for spectators to immediately understand the scoring criteria, such as technical points, composition points, and deductions. This can lead to questions such as, "This athlete looked better, but they got a low score," or "That athlete fell, but they got a high score." This reduces the enjoyment of watching the events and creates a sense of uncertainty regarding the scoring. [Means for solving the problem]
[0005] The present invention relates to a system including a means for receiving a video feed, a means for analyzing video data for each frame from the video feed and generating evaluation results in terms of technical scores, composition scores, and deductions, and a means for overlaying and displaying the evaluation results on the frames. This allows spectators to visually confirm the evaluation results for each performance in real time, making it easier for them to understand the scoring criteria. This can resolve spectator questions and improve the enjoyment of sports viewing. Furthermore, displaying the numerical analysis results in text increases the transparency of the evaluation results and improves spectator satisfaction.
[0006] "Video feed" refers to a continuous stream of video data captured in real time from a camera or video file.
[0007] "Video data" refers to the video information contained in each frame of a video feed, and analyzing this data can provide information including changes over time.
[0008] "Technical score" refers to the score used to evaluate the technical difficulty and perfection of sports performances and movements.
[0009] "Composition score" refers to the score used to evaluate the overall composition, balance, and beauty of a sports performance.
[0010] "Deduction" refers to the deduction of points when there is a mistake or an act that violates the rules in a sports performance or movement.
[0011] "Evaluation results" refers to comprehensive evaluation data regarding technical points, composition points, and deductions obtained by the analysis means.
[0012] "Overlay display" refers to the display of additional information (e.g., text or graphics) superimposed on the original video data.
[0013] "Real-time" refers to processing occurring immediately with virtually no time delay, meaning that the video feed is being analyzed and displayed simultaneously as it is received.
[0014] "Text display" refers to a method of displaying textual information on a screen, and refers to a means of visually presenting numerical data such as analytical results. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention relates to a system for displaying real-time scoring results during sports viewing. The system receives a video feed, analyzes it, generates technical scores, composition scores, and deductions, and visually displays the results. Specific embodiments of the system are described below.
[0037] System Overview
[0038] server
[0039] First, the server receives a video feed from a camera or a video file. The video feed contains footage of the athletes' performances and is processed frame by frame. The server initializes the video capture device and continuously receives the video data.
[0040] The server then extracts each frame from the received video feed and passes it to an AI evaluation module, which analyzes each frame and generates a technical score, composition score, and deductions for mistakes. This analysis result is obtained for each frame, allowing for real-time evaluation.
[0041] User terminal
[0042] The user's device receives the analysis results sent from the server and displays them as an overlay on the video feed frame. The numerical values for technical scores, composition scores, and deductions are displayed as text in the appropriate position within the frame. This allows the audience to visually check the evaluation score for each performance in real time, and intuitively understand the scoring criteria.
[0043] Program processing
[0044] server
[0045] 1. The server continues to receive the video feed from the video capture device.
[0046] 2. The server extracts each frame from the received video feed.
[0047] 3. The server passes each extracted frame to the AI evaluation module.
[0048] 4. The AI evaluation module analyzes each frame and generates an evaluation result including technical points, composition points, and deductions.
[0049] User terminal
[0050] 1. The user device receives the analysis results sent from the server.
[0051] 2. The user device overlays the analysis results on frames of the video feed.
[0052] 3. The user's device will display the numerical values for technical score, composition score, and deduction points for each frame in text at the appropriate position.
[0053] Specific examples
[0054] For example, in a figure skating competition, the server receives a frame from the video feed in which an athlete performs a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points to this jump, and a technical score of 7.8 points to the subsequent spin. The user device receives these evaluation results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can see this in real time and understand the performance evaluation on the spot. This improves the transparency and enjoyment of watching sports.
[0055] The processing flow will be explained below.
[0056] Step 1:
[0057] The server initializes the video capture device to receive the video feed from the camera or video file. A video capture object is created using cv2.VideoCapture(0), which starts capturing real-time video data.
[0058] Step 2:
[0059] The server reads the video data frame by frame from the received video feed, specifically by using the read method of the video capture object to extract the frame-by-frame images.
[0060] Step 3:
[0061] The server passes the acquired frames to the AI evaluation module, which receives the images frame by frame as preprocessing and extracts features related to the player's movements from the input.
[0062] Step 4:
[0063] The server receives the analysis results from the AI evaluation module, which calculates the technical score, composition score, and deduction score for each frame and returns the evaluation results including this data.
[0064] Step 5:
[0065] The server formats the evaluation results and converts them into a data format for sending to the user's device. Specifically, the evaluation results are converted into a flexible data format such as JSON.
[0066] Step 6:
[0067] The user terminal receives the analysis result data sent from the server, analyzes this data, and extracts the necessary information.
[0068] Step 7:
[0069] The user device overlays the analysis results on the video feed frame, using the cv2.putText method to display numerical information such as technical scores, composition scores, and deductions in the appropriate positions on the frame.
[0070] Step 8:
[0071] The user device displays the frame containing the overlay information on the screen in real time using the cv2.imshow method, allowing the audience to instantly see the evaluation results on the screen.
[0072] Specific examples
[0073] In a figure skating competition, the server receives the frame at the moment a triple axel jump is performed. The AI evaluation module analyzes this frame and assigns a technical score of 9.5 points for the jump. It also analyzes the frame for the subsequent spin and assigns a technical score of 7.8 points. These evaluation results are sent from the server to the user's device, which receives them. The user's device displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can view this in real time and instantly understand the evaluation of each performance. This process improves the transparency of scoring criteria and deepens understanding of sports spectators.
[0074] Example 1
[0075] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0076] Conventional sports viewing systems make it difficult for spectators to visually understand the technical and compositional scores of athletes in real time. Furthermore, the lack of transparency and speed in the evaluation process reduces the enjoyment of watching the game.
[0077] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0078] In this invention, the server includes means for receiving a video feed from a video capture device, means for extracting frames from the video feed and analyzing the video data to generate evaluation results regarding technical scores, composition scores, and deductions, means for overlaying the evaluation results on the frames in real time, and means for visually displaying the evaluation results in text format, thereby enabling spectators to visually check the evaluation scores in real time and enjoy fast and transparent evaluation.
[0079] "Video capture device" means equipment or software for receiving a video feed, including cameras and video streaming devices.
[0080] A "video feed" is a continuously distributed stream of moving image data, including footage of sporting matches and events.
[0081] "Frames" are the individual still images that make up a video feed and are played back in succession to form a moving image.
[0082] "Motion image data" is continuous image data consisting of multiple frames, and is used to analyze motion and movement.
[0083] "Technical score" is a technical evaluation score of an athlete's performance, based on the degree of perfection and difficulty of a particular movement.
[0084] "Composition score" is an evaluation score based on the composition and direction of the athlete's entire performance, taking into account creativity and expressiveness.
[0085] "Deductions" are points deducted from an athlete's technical and compositional scores for mistakes or incorrect movements during their performance.
[0086] The "evaluation result" is an analysis result including technical points, composition points, and deductions, and is generated for each frame.
[0087] "Overlay display" refers to the process of displaying the evaluation results overlaid on frames of the video feed, allowing the evaluation to be visually confirmed in real time.
[0088] "Text format" is a format in which information is displayed as text, and presents numerical values, evaluation points, etc. in a visually easy-to-understand manner.
[0089] A "deep learning evaluation module" is a module that uses artificial intelligence technology to analyze video data and evaluate technical and compositional aspects, and includes neural networks.
[0090] The present invention relates to a system for displaying real-time scoring results during sports viewing. The system receives video feeds from a video capture device, analyzes the video feeds, evaluates technical scores, composition scores, and deductions, and visually displays the evaluation results. Specific embodiments of the system are described below.
[0091] server
[0092] First, the server initializes a video capture device (e.g., HD camera, IP camera) as a video acquisition device and receives the video feed. The video feed contains footage of the athletes' performances, which is processed frame by frame. The server extracts the frames using an image processing library such as OpenCV.
[0093] Next, the server inputs each extracted frame into a deep learning evaluation module (e.g., a model using TensorFlow or PyTorch). This deep learning evaluation module analyzes each frame and evaluates technical points, composition points, and deductions. Analysis results are generated for each frame, allowing for real-time evaluation.
[0094] Specifically, the server receives a frame of a figure skating competition in which a skater performs a triple axel jump. This frame is preprocessed and sent to the evaluation module for technical evaluation. The evaluation module assigns a technical score of 9.5 points to the jump and 7.8 points to the subsequent spin.
[0095] User terminal
[0096] The user device receives the analysis results sent from the server. This data is sent and received using a real-time communication protocol such as WebSocket. Based on the received analysis results, the user device displays the evaluation results on the video feed frames. Specifically, numerical information such as technical scores, composition scores, and deductions is overlaid on the video frames using the HTML5 Canvas element and OpenGL.
[0097] Furthermore, the user's device will display the evaluation score in text format at the appropriate position for each frame. For example, text such as "Triple Axel: 9.5 points" or "Spin: 7.8 points" will be displayed in the bottom left or bottom right. This allows the audience to visually check the evaluation score in real time.
[0098] This system will improve the transparency of sports viewing and increase the enjoyment of watching, and will also enable spectators to intuitively understand the evaluation of performances in real time.
[0099] Prompt Sentence Examples
[0100] Below are some example prompts to input to the generative AI model:
[0101] "Describe an AI system that analyzes a video feed of figure skating performances and displays the technical and compositional scores for each performance in real time."
[0102] "Please tell us the specific technical details of a sports viewing system that displays the types of moves performed by athletes and their evaluation scores in real time."
[0103] The above is a specific embodiment for carrying out the present invention.
[0104] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0105] Step 1:
[0106] The server receives a video feed from a video capture device. The video capture device can be an HD camera or an IP camera. The input is the video feed from the video capture device, and the output is video image data divided into frames. This video feed is continuously captured and stored in the server's memory.
[0107] Step 2:
[0108] The server extracts individual frames from the video feed. It uses an image processing library such as OpenCV to acquire each frame. The input is the video feed received in step 1, and the output is image data for each frame. Specifically, it uses the OpenCV read function to extract the frames.
[0109] Step 3:
[0110] The server passes each extracted frame to a deep learning evaluation module. Using a deep learning framework such as TensorFlow or PyTorch, the module extracts and evaluates the frame's features. The input is the frame image extracted in step 2, and the output is the evaluation results, including technical scores, composition scores, and deductions. A convolutional neural network (CNN) is used for this evaluation to analyze the movement.
[0111] Step 4:
[0112] The deep learning evaluation module analyzes each frame and generates evaluation results in terms of technical points, composition points, and deductions. Specifically, it uses CNN to identify actions within the frame and calculate a score for each action. The input here is the frame image passed in step 3, and the output is the evaluation result for each frame.
[0113] Step 5:
[0114] The server sends the generated evaluation results to the user's device in real time using WebSocket. The input is the evaluation results generated in step 4, and the output is the analysis data sent to the user's device.
[0115] Step 6:
[0116] The user terminal receives the analysis results sent from the server. The input here is the evaluation results sent in step 5, and the output is the evaluation data stored in the memory on the user terminal.
[0117] Step 7:
[0118] The user device overlays the received analysis results on the video feed frame. Using the HTML5 Canvas element or OpenGL, the numerical information of the evaluation results is displayed overlaid on the video frame. The input is the evaluation results received in step 6, and the output is the overlaid video frame. Specific operations include displaying text in the appropriate position.
[0119] Step 8:
[0120] The user device displays the technical score, composition score, and deduction scores in text format at the appropriate position on the video frame. The input is the frame overlaid in step 7, and the output is real-time video that allows the audience to visually check the evaluation scores. For example, text information such as "Triple Axel: 9.5 points" or "Spin: 7.8 points" is displayed.
[0121] This series of processes allows spectators to visually check the evaluation scores for athletes' performances in real time, improving the transparency and enjoyment of watching sports.
[0122] (Application example 1)
[0123] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0124] With conventional sports viewing systems, it was difficult for spectators to check the evaluation of performances and plays in real time. In particular, there were limited means for them to visually understand the evaluation of technical points, composition points, and deductions. This reduced the transparency and enjoyment of the viewing experience, and prevented spectators from fully experiencing the appeal of sports. Another issue was the lack of systems that took into account use on mobile devices such as smartphones.
[0125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0126] In this invention, the server includes means for receiving a video feed, means for analyzing moving image data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for overlaying and displaying the evaluation results on the frames, means for receiving a video feed from a live streaming server, means for analyzing the video frames with an AI evaluation module and generating evaluation results, and means for overlaying and displaying the analysis results on the video feed in real time, thereby enabling spectators to visually understand the evaluation results of sports performances and plays being live-streamed in real time using their smartphones.
[0127] A "video feed" is a continuous stream of image data transmitted from a camera or other video capture device.
[0128] A "frame" is an individual still image in a video feed that is displayed in succession to form a moving image.
[0129] "Motion image data" refers to video data consisting of a series of frames obtained from a video feed.
[0130] "Technical points" are the points awarded to athletes in sports competitions based on their technical performance and skills.
[0131] "Composition points" are points given in sports competitions to evaluate the overall composition and artistry of an athlete's performance.
[0132] "Deductions" are points that are deducted in sports competitions when a player commits a mistake or acts in violation of the rules.
[0133] "AI Evaluation Module" means a software module that uses artificial intelligence technology to analyze video frames and generate technical scores, composition scores, and deduction scores.
[0134] A "live streaming server" is a server for delivering video feeds in real time over the Internet.
[0135] "Overlay display" is a technique for displaying additional information superimposed on a video frame.
[0136] "Real-time" means processed and displayed immediately, without delay or lag.
[0137] A "smartphone" is a portable information terminal that has mobile phone functions and can be used by installing additional applications.
[0138] "Analysis results" are the evaluation data of technical points, composition points, and deduction points generated by the AI evaluation module.
[0139] "Visual presentation" means displaying information in a way that is easy for the audience to see, and is a technique that allows information to be intuitively understood.
[0140] This invention relates to a system for displaying evaluation results in real time while watching sports. The system can receive a video feed, analyze it, generate evaluations of technical points, composition points, and deductions, and display the results visually.
[0141] System Overview
[0142] server
[0143] The server receives the video feed from the live streaming server and extracts video data from the video feed frame by frame. It initializes the video capture device and continuously receives video data. The server then extracts each frame from the received video feed and passes it to the AI evaluation module. The AI evaluation module analyzes each frame and generates technical scores, composition scores, and deductions for mistakes. The analysis results are obtained for each frame, allowing for real-time evaluation.
[0144] User terminal
[0145] The user's device receives the analysis results sent from the server and displays them as an overlay on the video feed frame. The numerical values for technical points, composition points, and deductions are displayed as text in the appropriate position within the frame. This allows spectators to visually check the evaluation scores for each performance and play in real time, and to intuitively understand the evaluation criteria.
[0146] Program processing
[0147] The implementation of this system requires several key software and hardware components: OpenCV is used to sample the video feed and extract frames, TensorFlow is used to generate scores for the AI evaluation module, and the phone's native UI framework (e.g., UIKit for iOS, Jetpack Compose for Android) is used to render the overlay display.
[0148] Specific hardware and software
[0149] Smartphone: iOS or Android device
[0150] Live Streaming Server: Streams video feeds using ffmpeg
[0151] AI evaluation module: Machine learning libraries such as TensorFlow
[0152] Frame analysis software: OpenCV
[0153] UI Framework: UIKit or Jetpack Compose
[0154] Specific examples
[0155] For example, in a figure skating competition, the server receives a frame from the video feed of the moment an athlete performs a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points for this jump and a technical score of 7.8 points for the subsequent spin. The user's device receives these evaluation results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can see this in real time and understand the performance evaluation on the spot. This improves the transparency and enjoyment of watching sports.
[0156] Prompt Sentence Examples
[0157] Create an application that displays real-time scoring results while watching sports. The app receives a video feed from a live streaming server and analyzes each frame with a TensorFlow model to generate and display technical scores, composition scores, and deductions. The software used is ffmpeg, OpenCV, and TensorFlow.
[0158] By implementing this invention, spectators can use their smartphones to visually understand the evaluation results of live-streamed sports performances and plays in real time.
[0159] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0160] Step 1:
[0161] The server receives a video feed from the live streaming server. It initializes a video capture device and continuously receives video data. The input is the video feed from the live streaming server, and the output is a stream of video frames. Specifically, the server initializes a video capture device and continuously processes the received video stream.
[0162] Step 2:
[0163] The server extracts each frame from the received video feed. It uses OpenCV to extract the frames and convert them into individual still image data. The input is the video feed, and the output is still image data for each frame. Specifically, it uses OpenCV functions to extract still images for each frame from the video stream.
[0164] Step 3:
[0165] The server passes each extracted frame to an AI evaluation module, which uses TensorFlow to analyze the frames and generate evaluations for technical, compositional, and deduction points. The input is still image data for each frame, and the output is the evaluation results for technical, compositional, and deduction points. Specifically, a pre-trained model is used to analyze the data for each frame and calculate the points.
[0166] Step 4:
[0167] The server transmits the evaluation results generated by the AI evaluation module to the user terminal. The evaluation results are transmitted in real time. The input is the evaluation result for each frame, and the output is the evaluation data transmitted over the network. Specifically, the server converts the evaluation results into an appropriate format and transmits them over the network.
[0168] Step 5:
[0169] The user device overlays the analysis results received from the server on the video feed frame. The input is the evaluation results sent from the server, and the output is the overlaid video frame. Specifically, the device's native UI framework is used to display the evaluation results in text on the frame in real time.
[0170] Step 6:
[0171] The user checks the evaluation results in real time using a smartphone. The input is the video feed displayed on the user's device, and the output is the user's visual understanding of the information. Specifically, the user checks the performances and plays in the video stream along with the evaluation scores overlaid in real time.
[0172] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0173] The present invention relates to a system that displays real-time scoring results during sports viewing and customizes the display content by recognizing the user's emotions. The system receives a video feed, analyzes it, generates technical scores, composition scores, and deductions, and visually displays the results. It also incorporates an emotion engine that recognizes the user's emotions. Specific embodiments of the system are described below.
[0174] System Overview
[0175] server
[0176] First, the server receives a video feed from a camera or a video file. The video feed contains footage of the athletes' performances and is processed frame by frame. The server initializes the video capture device and continuously receives the video data.
[0177] The server then extracts each frame from the received video feed and passes it to an AI evaluation module, which analyzes each frame and generates a technical score, composition score, and deductions for mistakes. This analysis result is obtained for each frame, allowing for real-time evaluation.
[0178] Furthermore, the server is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's facial expressions and voice to identify emotions such as enjoyment, surprise, or dissatisfaction. Based on the recognized emotion, the server customizes the display of the evaluation results. For example, if the user shows interest, it can provide additional detailed information.
[0179] User terminal
[0180] The user's device receives the analysis results sent from the server as well as information from the emotion engine. Based on this, an overlay display is performed on the video feed frame. As text, numerical values for technical points, composition points, and deductions are displayed in appropriate positions within the frame. Furthermore, a UI (user interface) can be constructed according to the user's emotions, and the displayed content can be dynamically changed.
[0181] Program processing
[0182] server
[0183] 1. The server continues to receive the video feed from the video capture device.
[0184] 2. The server extracts each frame from the received video feed.
[0185] 3. The server passes each extracted frame to the AI evaluation module.
[0186] 4. The AI evaluation module analyzes each frame and generates an evaluation result including technical points, composition points, and deductions.
[0187] 5. The server uses an emotion engine to recognize emotions from the user's face and voice.
[0188] 6. Customize the display of rating results based on the emotions recognized by the emotion engine.
[0189] User terminal
[0190] 1. The user device receives the analysis results and emotion engine information sent from the server.
[0191] 2. The user device overlays the analysis results and emotion data onto frames of the video feed.
[0192] 3. The user's device will display the numerical values for technical score, composition score, and deduction points for each frame in text at the appropriate position.
[0193] 4. Dynamically change the UI based on the user's emotions, for example highlighting important information if the user is surprised.
[0194] Specific examples
[0195] For example, in a figure skating competition, the server receives a frame from the video feed of a skater performing a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points for the jump and a technical score of 7.8 points for the subsequent spin. These evaluation results are then sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice and recognizes their surprise. The user's device receives the results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." To further emphasize their surprise, the user's device also displays a comment such as "Great jump!" Spectators can view this in real time and receive interactive feedback based on the evaluation of each performance and their own emotions. This process increases the transparency of scoring criteria and further enhances the enjoyment of watching sports.
[0196] The processing flow will be explained below.
[0197] Step 1:
[0198] The server initializes a video capture device to receive the video feed from a camera or a video file. Specifically, it uses cv2.VideoCapture(0) to get real-time video data from the device.
[0199] Step 2:
[0200] The server extracts individual frames from the received video feed, using the read method of the video capture object to retrieve the image data for each frame.
[0201] Step 3:
[0202] The server passes each extracted frame to an AI evaluation module, which uses a pre-trained model to analyze the content of each frame and calculate technical and compositional scores, as well as deductions.
[0203] Step 4:
[0204] The server receives the analysis results from the AI evaluation module and converts them into a data format (e.g., JSON format), which is then prepared for transmission to the user's device.
[0205] Step 5:
[0206] The server uses an emotion engine to analyze the user's facial expressions and voice in real time to recognize their emotions. For example, it analyzes camera footage and microphone input to determine whether the user is enjoying, surprised, or dissatisfied.
[0207] Step 6:
[0208] The server customizes the displayed evaluation results based on the user's emotions. For example, if the user is surprised, it highlights specific technical or structural features.
[0209] Step 7:
[0210] The server then sends the customized evaluation results to the user's device, using a network to transfer data in real time.
[0211] Step 8:
[0212] The user device receives the analysis result data and emotion data sent from the server, and prepares an overlay display on the frame of the video feed based on this data.
[0213] Step 9:
[0214] The user's device overlays the analysis results on the frame. Specifically, the cv2.putText method is used to display the numerical information on technical points, composition points, and deductions in text format in the appropriate location.
[0215] Step 10:
[0216] The user device dynamically changes the UI based on the user's emotions. For example, if it detects that the user is surprised, it will display a comment such as "Great jump!"
[0217] Step 11:
[0218] The user device displays frames containing overlay information on the screen in real time, allowing the audience to see the evaluation results and interactive feedback in real time.
[0219] Specific examples
[0220] In a figure skating competition, the server receives a frame from the video feed of the moment an athlete performs a triple axel jump. The AI evaluation module analyzes this frame and assigns a technical score of 9.5 points for the jump. It also analyzes the frame for the spin and assigns a technical score of 7.8 points. These evaluation results are sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice to recognize surprise. The user's device receives this and overlays text such as "Triple axel: 9.5 points" and "Spin: 7.8 points" on the frame. It also displays a comment such as "Great jump!" to emphasize surprise. This allows spectators to receive interactive feedback in real time based on the evaluation of each performance and their own emotions, improving the transparency and enjoyment of sports viewing.
[0221] Example 2
[0222] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0223] Conventional sports viewing systems have the problem that it is difficult to visually display the performance evaluation of athletes in real time, and they do not provide interactive feedback according to the user's emotions. As a result, spectators cannot immediately grasp the detailed evaluation results during the performance, which limits the viewing experience.
[0224] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0225] In this invention, the server includes means for receiving a video feed, means for analyzing video data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for recognizing a user's emotions and customizing the display content, and means for overlaying the evaluation results on the frames, thereby making it possible to accurately evaluate a player's performance in real time, display the evaluation results, and provide interactive feedback according to the user's emotions.
[0226] "Video Feed" refers to a continuous stream of real-time or recorded video data.
[0227] "Motion image data" refers to digital data that contains image information that changes over time, and is typically made up of a collection of individual still images called frames.
[0228] "Technical score" is a numerical value that indicates the evaluation of the technical elements of an athlete's performance, including the accuracy and difficulty of technical movements such as jumps and spins.
[0229] "Composition score" is a numerical value that indicates an evaluation of the composition and artistic quality of a performance, and includes, for example, expressiveness and the flow of the performance.
[0230] "Deductions" are points that are subtracted from the technical or composition score due to a mistake or violation of rules.
[0231] "User emotion" refers to the subjective emotional state, such as enjoyment, surprise, or dissatisfaction, that a spectator or user feels while watching a sporting event.
[0232] "Customizing the display content" refers to dynamically changing the information and interface content displayed in response to the user's emotions.
[0233] "Overlay display" refers to a method of displaying evaluation results or additional information overlaid on the video feed.
[0234] This invention relates to a system that displays the scores of sports games in real time and customizes the display content by recognizing the emotions of users. The system is composed of a server and a user terminal.
[0235] First, the server initializes the video capture device and receives the athletes' performances in real time. Specifically, it uses OpenCV to acquire video data from IP cameras and USB cameras. The received video feed is extracted frame by frame and passed to the AI evaluation module. The AI evaluation module uses TensorFlow and PyTorch to analyze each frame and generate technical scores, composition scores, and deductions. This makes it possible to evaluate the athletes' performances in real time.
[0236] The server is also equipped with an emotion engine that recognizes the user's emotions. Using services such as Google Cloud Natural Language API and IBM Watson, it analyzes the user's facial expressions and voice to identify emotions such as whether the user is enjoying, surprised, or dissatisfied. After identifying the emotion, the server customizes the display of the evaluation results. For example, if the user is surprised, it can add a comment such as "Great jump!"
[0237] Next, the user device receives the analysis results and emotion engine information sent from the server. <canvas>Using JavaScript and JavaScript, the analysis results are overlaid on frames of the video feed. By visually displaying numerical values for technical, compositional, and deduction scores, the audience can see a detailed evaluation of each performance in real time. The UI also dynamically changes based on the user's emotions, highlighting important information as needed.
[0238] Take a figure skating competition as an example. The server receives the moment of a triple axel jump from a video feed, and the AI evaluation module assigns a technical score of 9.5 points for this jump and 7.8 points for the following spin. These evaluation results are sent to the user's device in real time, and the emotion engine recognizes that the user is surprised. The user's device receives this and displays a text overlay saying "Triple axel: 9.5 points" and "Spin: 7.8 points," along with the comment "Great jump!"
[0239] Example prompt sentence:
[0240] "Analyze a figure skater's triple axel jump in real time and rate it for technical and compositional points. Highlight results that amaze the user."
[0241] As described above, this invention allows users to check player evaluations in real time and receive interactive feedback based on their own emotions. This system will further enhance the enjoyment of watching sports.
[0242] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0243] Step 1:
[0244] The server initializes the video capture device and receives the athletes' performances in real time. Specifically, the server acquires video data using OpenCV's cv2.VideoCapture. The input is video data from the video capture device, and the output is video data divided into individual frames.
[0245] Step 2:
[0246] The server extracts frames one by one from the received video feed. Specifically, the server reads frames continuously using the OpenCV read() method and adds them to a list. The input is continuous video data from the video capture device, and the output is video data frame by frame.
[0247] Step 3:
[0248] The server passes each extracted frame to the AI evaluation module. Specifically, the server inputs the frame data into the AI evaluation module using TensorFlow or PyTorch. The input is video data for each frame, and the output is analyzed evaluation data (technical score, composition score, and deductions).
[0249] Step 4:
[0250] The AI evaluation module analyzes each frame and generates technical scores, composition scores, and deductions. Specifically, the AI model analyzes specific movements within a frame and outputs evaluation data such as "jump: 9.5 points" or "spin: 7.8 points." The input is the video data for each frame, and the output is evaluation data for each performance.
[0251] Step 5:
[0252] The server uses an emotion engine to recognize emotions from the user's face and voice. Specifically, the server analyzes the user's facial expression and voice data using Google Cloud Natural Language API and IBM Watson. The input is the user's facial expression and voice data, and the output is the user's emotional state (enjoyed, surprised, dissatisfied, etc.).
[0253] Step 6:
[0254] The display content is customized based on the emotions recognized by the emotion engine. Specifically, the server generates additional information or comments to attract the user's attention based on the analysis results. For example, if the user is surprised, it adds a comment such as "Great jump!" The input is the user's emotional state and evaluation data for each frame, and the output is customized display data.
[0255] Step 7:
[0256] The user device receives the analysis results and emotion engine information sent from the server. Specifically, the user device communicates with the server using WebSocket or RESTful API to receive data. The input is the analysis results and emotion information sent from the server, and the output is data stored on the user device.
[0257] Step 8:
[0258] The user device overlays the analysis results and emotion data on the video feed frame. <canvas>It uses elements and JavaScript to overlay analytics data on a video feed. The input is analytics results and emotion data, and the output is a visual display on the user's device.
[0259] Step 9:
[0260] The user's device displays the technical score, composition score, and deduction score values for each frame in the appropriate position. Specifically, it performs coordinate calculations to display the analysis data as text in the appropriate position within the frame. The input is the evaluation data for each frame, and the output is a text overlay display for each frame.
[0261] Step 10:
[0262] The user device dynamically changes the UI according to the user's emotions. Specifically, it uses CSS and JavaScript to change UI elements and highlight important information. For example, if the user is surprised, it changes the text color or highlights a specific comment. The input is the user's emotional data and analysis results, and the output is a dynamically changed UI display.
[0263] (Application example 2)
[0264] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0265] In conventional sports viewing systems, it was difficult for spectators to know the evaluation results of athletes' technical and compositional scores in real time, and the evaluation criteria lacked transparency. Furthermore, there was no way to provide a more interactive and engaging viewing experience by incorporating spectators' emotions and excitement. As a result, it was difficult to increase spectator satisfaction.
[0266] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0267] In this invention, the server includes means for receiving a video feed, means for analyzing video data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for overlaying the evaluation results on the frames, and means for analyzing a user's emotions and customizing the display content of the evaluation results. This allows viewers to know the evaluation results, such as technical scores and composition scores, in real time when watching a sports game, and further provides an interactive viewing experience that responds to the viewers' emotions.
[0268] A "video feed" is a continuous video signal coming from a camera or other video capture device.
[0269] A "frame" refers to an individual still image within a video feed, with a series of frames forming a moving image.
[0270] "Video data" refers to the digital video data that comprises each frame in a video feed.
[0271] "Technical score" is a numerical evaluation of an athlete's technical performance in a sporting event.
[0272] "Composition points" are a numerical evaluation of the composition and expressiveness of a performance in a sporting event.
[0273] "Deduction points" refers to the evaluation of a sporting event in which negative points are given for technical errors or violations of rules.
[0274] "Evaluation results" refers to the overall score including technical points, composition points, and deductions.
[0275] "Overlay display" refers to a technique for displaying additional information overlaid on top of a video frame.
[0276] "User's emotion" refers to the subjective emotional state that the user feels while watching a sporting event, and is analyzed from facial expressions and voice.
[0277] "Emotion engine" refers to a software engine that recognizes and analyzes emotions from the user's face and voice.
[0278] "UI" stands for user interface and refers to the screen layout and design that allows users to interact with systems and applications.
[0279] "Real-time" refers to information and data processing occurring in real time, with results reflected almost immediately.
[0280] A system for implementing this invention consists of a server and a user terminal. The server receives a video feed and analyzes the video data frame by frame to generate evaluation results in terms of technical points, composition points, and deductions. It also analyzes the user's emotions and customizes the display of the evaluation results. The user terminal displays the evaluation results as an overlay, visually presenting them to the user.
[0281] Specifically, the server receives live feeds from cameras and video capture devices. This live feed contains the athletes' performances and analyzes them frame by frame. An AI evaluation module using TensorFlow is used for the analysis, and evaluation results are generated in real time, including technical scores, compositional scores, and deductions. These evaluation results are then sent to the user's device.
[0282] The server is equipped with an emotion engine that recognizes emotions from viewers' faces and voices. This emotion engine uses VaderSentiment and TensorFlow emotion recognition models. Viewers' facial expressions and voice data are analyzed using the dlib library to grasp the user's emotional state in real time. As a result, the display of the evaluation results is customized, such as highlighting important information if the user is surprised.
[0283] The user's device receives the analysis results and emotion data sent from the server. A Python-based GUI application is installed on the user's device, which overlays the live feed frames with information on technical and compositional points, as well as deductions. Furthermore, the application dynamically displays comments and supplementary information based on the user's emotions.
[0284] As a concrete example, in a figure skating competition, the server analyzes the frame at the moment when the skater performs a triple axel jump. The AI evaluation module generates a technical score of 9.5 points for this jump, and a technical score of 7.8 points for the subsequent spin. These evaluation results are sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice and recognizes that the user is surprised. The user's device receives this and overlays the frame with text such as "Triple axel: 9.5 points" and "Spin: 7.8 points," along with additional comments such as "Great jump!"
[0285] An example prompt is:
[0286] "Design a system that displays technical and compositional scores in real time based on viewer emotions during live-streamed sporting events. Utilize an AI evaluation module and emotion recognition engine to dynamically display cheering comments, such as 'Great jump!', when a viewer expresses surprise."
[0287] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0288] Step 1:
[0289] The server receives the video feed from the camera or video capture device. Here, the input is a live video feed, and each frame of the video feed is continuously sent to the server. This step captures the athlete's performance in real time. The output is a continuous video signal.
[0290] Step 2:
[0291] The server extracts each frame from the received video feed. The input is the video feed, and data is extracted frame by frame. Here, the actual processing is performed to extract the video data obtained from the video capture device one frame at a time. The output is frame data, which are individual still images.
[0292] Step 3:
[0293] The server passes the extracted frame data to the AI evaluation module. The input is the frame data, which is analyzed frame by frame using TensorFlow. The module calculates the technical score, composition score, and deduction points, and outputs the evaluation results as numerical data.
[0294] Step 4:
[0295] The server is equipped with an emotion engine that analyzes emotions from the user's facial expressions and voice. The input is the user's facial expression data and voice data, which are analyzed using the dlib library and VaderSentiment. In this step, specific processing is performed to identify the user's emotional state, such as whether they are surprised, amused, or frustrated. The output is the user's emotional state data.
[0296] Step 5:
[0297] The server combines the evaluation results and the user's emotional state to customize the content of the evaluation results. The input is the numerical evaluation results and emotional state data, and the display content is dynamically changed based on these. In this step, specific operations are performed, such as adding a comment such as "Great jump!" if the user is surprised. The output is customized evaluation result data.
[0298] Step 6:
[0299] The server sends the customized evaluation result data to the user terminal. The input is the customized evaluation result data, which is sent to the user terminal via the network. In this step, the specific operation of data transfer is performed. The output is the evaluation result data sent to the user terminal.
[0300] Step 7:
[0301] The user device uses the received evaluation result data and emotion data to overlay a frame of the video feed. The input is the customized evaluation result data and emotion data, and specific processing is performed to display text and comments based on these. The output is a visual display result.
[0302] Step 8:
[0303] The user device dynamically changes the UI according to the user's emotional response. The input is real-time emotional data, and specific actions are taken to change the interface and presented information based on this. The output is an updated UI.
[0304] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0305] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0306] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0307] [Second embodiment]
[0308] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0309] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0310] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0311] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0312] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0313] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0314] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0315] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0316] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0317] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0318] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0319] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0320] The present invention relates to a system for displaying real-time scoring results during sports viewing. The system receives a video feed, analyzes it, generates technical scores, composition scores, and deductions, and visually displays the results. Specific embodiments of the system are described below.
[0321] System Overview
[0322] server
[0323] First, the server receives a video feed from a camera or a video file. The video feed contains footage of the athletes' performances and is processed frame by frame. The server initializes the video capture device and continuously receives the video data.
[0324] The server then extracts each frame from the received video feed and passes it to an AI evaluation module, which analyzes each frame and generates a technical score, composition score, and deductions for mistakes. This analysis result is obtained for each frame, allowing for real-time evaluation.
[0325] User terminal
[0326] The user's device receives the analysis results sent from the server and displays them as an overlay on the video feed frame. The numerical values for technical scores, composition scores, and deductions are displayed as text in the appropriate position within the frame. This allows the audience to visually check the evaluation score for each performance in real time, and intuitively understand the scoring criteria.
[0327] Program processing
[0328] server
[0329] 1. The server continues to receive the video feed from the video capture device.
[0330] 2. The server extracts each frame from the received video feed.
[0331] 3. The server passes each extracted frame to the AI evaluation module.
[0332] 4. The AI evaluation module analyzes each frame and generates an evaluation result including technical points, composition points, and deductions.
[0333] User terminal
[0334] 1. The user device receives the analysis results sent from the server.
[0335] 2. The user device overlays the analysis results on frames of the video feed.
[0336] 3. The user's device will display the numerical values for technical score, composition score, and deduction points for each frame in text at the appropriate position.
[0337] Specific examples
[0338] For example, in a figure skating competition, the server receives a frame from the video feed in which an athlete performs a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points to this jump, and a technical score of 7.8 points to the subsequent spin. The user device receives these evaluation results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can see this in real time and understand the performance evaluation on the spot. This improves the transparency and enjoyment of watching sports.
[0339] The processing flow will be explained below.
[0340] Step 1:
[0341] The server initializes the video capture device to receive the video feed from the camera or video file. A video capture object is created using cv2.VideoCapture(0), which starts capturing real-time video data.
[0342] Step 2:
[0343] The server reads the video data frame by frame from the received video feed, specifically by using the read method of the video capture object to extract the frame-by-frame images.
[0344] Step 3:
[0345] The server passes the acquired frames to the AI evaluation module, which receives the images frame by frame as preprocessing and extracts features related to the player's movements from the input.
[0346] Step 4:
[0347] The server receives the analysis results from the AI evaluation module, which calculates the technical score, composition score, and deduction score for each frame and returns the evaluation results including this data.
[0348] Step 5:
[0349] The server formats the evaluation results and converts them into a data format for sending to the user's device. Specifically, the evaluation results are converted into a flexible data format such as JSON.
[0350] Step 6:
[0351] The user terminal receives the analysis result data sent from the server, analyzes this data, and extracts the necessary information.
[0352] Step 7:
[0353] The user device overlays the analysis results on the video feed frame, using the cv2.putText method to display numerical information such as technical scores, composition scores, and deductions in the appropriate positions on the frame.
[0354] Step 8:
[0355] The user device displays the frame containing the overlay information on the screen in real time using the cv2.imshow method, allowing the audience to instantly see the evaluation results on the screen.
[0356] Specific examples
[0357] In a figure skating competition, the server receives the frame at the moment a triple axel jump is performed. The AI evaluation module analyzes this frame and assigns a technical score of 9.5 points for the jump. It also analyzes the frame for the subsequent spin and assigns a technical score of 7.8 points. These evaluation results are sent from the server to the user's device, which receives them. The user's device displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can view this in real time and instantly understand the evaluation of each performance. This process improves the transparency of scoring criteria and deepens understanding of sports spectators.
[0358] Example 1
[0359] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0360] Conventional sports viewing systems make it difficult for spectators to visually understand the technical and compositional scores of athletes in real time. Furthermore, the lack of transparency and speed in the evaluation process reduces the enjoyment of watching the game.
[0361] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0362] In this invention, the server includes means for receiving a video feed from a video capture device, means for extracting frames from the video feed and analyzing the video data to generate evaluation results regarding technical scores, composition scores, and deductions, means for overlaying the evaluation results on the frames in real time, and means for visually displaying the evaluation results in text format, thereby enabling spectators to visually check the evaluation scores in real time and enjoy fast and transparent evaluation.
[0363] "Video capture device" means equipment or software for receiving a video feed, including cameras and video streaming devices.
[0364] A "video feed" is a continuously distributed stream of moving image data, including footage of sporting matches and events.
[0365] "Frames" are the individual still images that make up a video feed and are played back in succession to form a moving image.
[0366] "Motion image data" is continuous image data consisting of multiple frames, and is used to analyze motion and movement.
[0367] "Technical score" is a technical evaluation score of an athlete's performance, based on the degree of perfection and difficulty of a particular movement.
[0368] "Composition score" is an evaluation score based on the composition and direction of the athlete's entire performance, taking into account creativity and expressiveness.
[0369] "Deductions" are points deducted from an athlete's technical and compositional scores for mistakes or incorrect movements during their performance.
[0370] The "evaluation result" is an analysis result including technical points, composition points, and deductions, and is generated for each frame.
[0371] "Overlay display" refers to the process of displaying the evaluation results overlaid on frames of the video feed, allowing the evaluation to be visually confirmed in real time.
[0372] "Text format" is a format in which information is displayed as text, and presents numerical values, evaluation points, etc. in a visually easy-to-understand manner.
[0373] A "deep learning evaluation module" is a module that uses artificial intelligence technology to analyze video data and evaluate technical and compositional aspects, and includes neural networks.
[0374] The present invention relates to a system for displaying real-time scoring results during sports viewing. The system receives video feeds from a video capture device, analyzes the video feeds, evaluates technical scores, composition scores, and deductions, and visually displays the evaluation results. Specific embodiments of the system are described below.
[0375] server
[0376] First, the server initializes a video capture device (e.g., HD camera, IP camera) as a video acquisition device and receives the video feed. The video feed contains footage of the athletes' performances, which is processed frame by frame. The server extracts the frames using an image processing library such as OpenCV.
[0377] Next, the server inputs each extracted frame into a deep learning evaluation module (e.g., a model using TensorFlow or PyTorch). This deep learning evaluation module analyzes each frame and evaluates technical points, composition points, and deductions. Analysis results are generated for each frame, allowing for real-time evaluation.
[0378] Specifically, the server receives a frame of a figure skating competition in which a skater performs a triple axel jump. This frame is preprocessed and sent to the evaluation module for technical evaluation. The evaluation module assigns a technical score of 9.5 points to the jump and 7.8 points to the subsequent spin.
[0379] User terminal
[0380] The user device receives the analysis results sent from the server. This data is sent and received using a real-time communication protocol such as WebSocket. Based on the received analysis results, the user device displays the evaluation results on the video feed frames. Specifically, numerical information such as technical scores, composition scores, and deductions is overlaid on the video frames using the HTML5 Canvas element and OpenGL.
[0381] Furthermore, the user's device will display the evaluation score in text format at the appropriate position for each frame. For example, text such as "Triple Axel: 9.5 points" or "Spin: 7.8 points" will be displayed in the bottom left or bottom right. This allows the audience to visually check the evaluation score in real time.
[0382] This system will improve the transparency of sports viewing and increase the enjoyment of watching, and will also enable spectators to intuitively understand the evaluation of performances in real time.
[0383] Prompt Sentence Examples
[0384] Below are some example prompts to input to the generative AI model:
[0385] "Describe an AI system that analyzes a video feed of figure skating performances and displays the technical and compositional scores for each performance in real time."
[0386] "Please tell us the specific technical details of a sports viewing system that displays the types of moves performed by athletes and their evaluation scores in real time."
[0387] The above is a specific embodiment for carrying out the present invention.
[0388] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0389] Step 1:
[0390] The server receives a video feed from a video capture device. The video capture device can be an HD camera or an IP camera. The input is the video feed from the video capture device, and the output is video image data divided into frames. This video feed is continuously captured and stored in the server's memory.
[0391] Step 2:
[0392] The server extracts individual frames from the video feed. It uses an image processing library such as OpenCV to acquire each frame. The input is the video feed received in step 1, and the output is image data for each frame. Specifically, it uses the OpenCV read function to extract the frames.
[0393] Step 3:
[0394] The server passes each extracted frame to a deep learning evaluation module. Using a deep learning framework such as TensorFlow or PyTorch, the module extracts and evaluates the frame's features. The input is the frame image extracted in step 2, and the output is the evaluation results, including technical scores, composition scores, and deductions. A convolutional neural network (CNN) is used for this evaluation to analyze the movement.
[0395] Step 4:
[0396] The deep learning evaluation module analyzes each frame and generates evaluation results in terms of technical points, composition points, and deductions. Specifically, it uses CNN to identify actions within the frame and calculate a score for each action. The input here is the frame image passed in step 3, and the output is the evaluation result for each frame.
[0397] Step 5:
[0398] The server sends the generated evaluation results to the user's device in real time using WebSocket. The input is the evaluation results generated in step 4, and the output is the analysis data sent to the user's device.
[0399] Step 6:
[0400] The user terminal receives the analysis results sent from the server. The input here is the evaluation results sent in step 5, and the output is the evaluation data stored in the memory on the user terminal.
[0401] Step 7:
[0402] The user device overlays the received analysis results on the video feed frame. Using the HTML5 Canvas element or OpenGL, the numerical information of the evaluation results is displayed overlaid on the video frame. The input is the evaluation results received in step 6, and the output is the overlaid video frame. Specific operations include displaying text in the appropriate position.
[0403] Step 8:
[0404] The user device displays the technical score, composition score, and deduction scores in text format at the appropriate position on the video frame. The input is the frame overlaid in step 7, and the output is real-time video that allows the audience to visually check the evaluation scores. For example, text information such as "Triple Axel: 9.5 points" or "Spin: 7.8 points" is displayed.
[0405] This series of processes allows spectators to visually check the evaluation scores for athletes' performances in real time, improving the transparency and enjoyment of watching sports.
[0406] (Application example 1)
[0407] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0408] With conventional sports viewing systems, it was difficult for spectators to check the evaluation of performances and plays in real time. In particular, there were limited means for them to visually understand the evaluation of technical points, composition points, and deductions. This reduced the transparency and enjoyment of the viewing experience, and prevented spectators from fully experiencing the appeal of sports. Another issue was the lack of systems that took into account use on mobile devices such as smartphones.
[0409] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0410] In this invention, the server includes means for receiving a video feed, means for analyzing moving image data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for overlaying and displaying the evaluation results on the frames, means for receiving a video feed from a live streaming server, means for analyzing the video frames with an AI evaluation module and generating evaluation results, and means for overlaying and displaying the analysis results on the video feed in real time, thereby enabling spectators to visually understand the evaluation results of sports performances and plays being live-streamed in real time using their smartphones.
[0411] A "video feed" is a continuous stream of image data transmitted from a camera or other video capture device.
[0412] A "frame" is an individual still image in a video feed that is displayed in succession to form a moving image.
[0413] "Motion image data" refers to video data consisting of a series of frames obtained from a video feed.
[0414] "Technical points" are the points awarded to athletes in sports competitions based on their technical performance and skills.
[0415] "Composition points" are points given in sports competitions to evaluate the overall composition and artistry of an athlete's performance.
[0416] "Deductions" are points that are deducted in sports competitions when a player commits a mistake or acts in violation of the rules.
[0417] "AI Evaluation Module" means a software module that uses artificial intelligence technology to analyze video frames and generate technical scores, composition scores, and deduction scores.
[0418] A "live streaming server" is a server for delivering video feeds in real time over the Internet.
[0419] "Overlay display" is a technique for displaying additional information superimposed on a video frame.
[0420] "Real-time" means processed and displayed immediately, without delay or lag.
[0421] A "smartphone" is a portable information terminal that has mobile phone functions and can be used by installing additional applications.
[0422] "Analysis results" are the evaluation data of technical points, composition points, and deduction points generated by the AI evaluation module.
[0423] "Visual presentation" means displaying information in a way that is easy for the audience to see, and is a technique that allows information to be intuitively understood.
[0424] This invention relates to a system for displaying evaluation results in real time while watching sports. The system can receive a video feed, analyze it, generate evaluations of technical points, composition points, and deductions, and display the results visually.
[0425] System Overview
[0426] server
[0427] The server receives the video feed from the live streaming server and extracts video data from the video feed frame by frame. It initializes the video capture device and continuously receives video data. The server then extracts each frame from the received video feed and passes it to the AI evaluation module. The AI evaluation module analyzes each frame and generates technical scores, composition scores, and deductions for mistakes. The analysis results are obtained for each frame, allowing for real-time evaluation.
[0428] User terminal
[0429] The user's device receives the analysis results sent from the server and displays them as an overlay on the video feed frame. The numerical values for technical points, composition points, and deductions are displayed as text in the appropriate position within the frame. This allows spectators to visually check the evaluation scores for each performance and play in real time, and to intuitively understand the evaluation criteria.
[0430] Program processing
[0431] The implementation of this system requires several key software and hardware components: OpenCV is used to sample the video feed and extract frames, TensorFlow is used to generate scores for the AI evaluation module, and the phone's native UI framework (e.g., UIKit for iOS, Jetpack Compose for Android) is used to render the overlay display.
[0432] Specific hardware and software
[0433] Smartphone: iOS or Android device
[0434] Live Streaming Server: Streams video feeds using ffmpeg
[0435] AI evaluation module: Machine learning libraries such as TensorFlow
[0436] Frame analysis software: OpenCV
[0437] UI Framework: UIKit or Jetpack Compose
[0438] Specific examples
[0439] For example, in a figure skating competition, the server receives a frame from the video feed of the moment an athlete performs a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points for this jump and a technical score of 7.8 points for the subsequent spin. The user's device receives these evaluation results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can see this in real time and understand the performance evaluation on the spot. This improves the transparency and enjoyment of watching sports.
[0440] Prompt Sentence Examples
[0441] Create an application that displays real-time scoring results while watching sports. The app receives a video feed from a live streaming server and analyzes each frame with a TensorFlow model to generate and display technical scores, composition scores, and deductions. The software used is ffmpeg, OpenCV, and TensorFlow.
[0442] By implementing this invention, spectators can use their smartphones to visually understand the evaluation results of live-streamed sports performances and plays in real time.
[0443] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0444] Step 1:
[0445] The server receives a video feed from the live streaming server. It initializes a video capture device and continuously receives video data. The input is the video feed from the live streaming server, and the output is a stream of video frames. Specifically, the server initializes a video capture device and continuously processes the received video stream.
[0446] Step 2:
[0447] The server extracts each frame from the received video feed. It uses OpenCV to extract the frames and convert them into individual still image data. The input is the video feed, and the output is still image data for each frame. Specifically, it uses OpenCV functions to extract still images for each frame from the video stream.
[0448] Step 3:
[0449] The server passes each extracted frame to an AI evaluation module, which uses TensorFlow to analyze the frames and generate evaluations for technical, compositional, and deduction points. The input is still image data for each frame, and the output is the evaluation results for technical, compositional, and deduction points. Specifically, a pre-trained model is used to analyze the data for each frame and calculate the points.
[0450] Step 4:
[0451] The server transmits the evaluation results generated by the AI evaluation module to the user terminal. The evaluation results are transmitted in real time. The input is the evaluation result for each frame, and the output is the evaluation data transmitted over the network. Specifically, the server converts the evaluation results into an appropriate format and transmits them over the network.
[0452] Step 5:
[0453] The user device overlays the analysis results received from the server on the video feed frame. The input is the evaluation results sent from the server, and the output is the overlaid video frame. Specifically, the device's native UI framework is used to display the evaluation results in text on the frame in real time.
[0454] Step 6:
[0455] The user checks the evaluation results in real time using a smartphone. The input is the video feed displayed on the user's device, and the output is the user's visual understanding of the information. Specifically, the user checks the performances and plays in the video stream along with the evaluation scores overlaid in real time.
[0456] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0457] The present invention relates to a system that displays real-time scoring results during sports viewing and customizes the display content by recognizing the user's emotions. The system receives a video feed, analyzes it, generates technical scores, composition scores, and deductions, and visually displays the results. It also incorporates an emotion engine that recognizes the user's emotions. Specific embodiments of the system are described below.
[0458] System Overview
[0459] server
[0460] First, the server receives a video feed from a camera or a video file. The video feed contains footage of the athletes' performances and is processed frame by frame. The server initializes the video capture device and continuously receives the video data.
[0461] The server then extracts each frame from the received video feed and passes it to an AI evaluation module, which analyzes each frame and generates a technical score, composition score, and deductions for mistakes. This analysis result is obtained for each frame, allowing for real-time evaluation.
[0462] Furthermore, the server is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's facial expressions and voice to identify emotions such as enjoyment, surprise, or dissatisfaction. Based on the recognized emotion, the server customizes the display of the evaluation results. For example, if the user shows interest, it can provide additional detailed information.
[0463] User terminal
[0464] The user's device receives the analysis results sent from the server as well as information from the emotion engine. Based on this, an overlay display is performed on the video feed frame. As text, numerical values for technical points, composition points, and deductions are displayed in appropriate positions within the frame. Furthermore, a UI (user interface) can be constructed according to the user's emotions, and the displayed content can be dynamically changed.
[0465] Program processing
[0466] server
[0467] 1. The server continues to receive the video feed from the video capture device.
[0468] 2. The server extracts each frame from the received video feed.
[0469] 3. The server passes each extracted frame to the AI evaluation module.
[0470] 4. The AI evaluation module analyzes each frame and generates an evaluation result including technical points, composition points, and deductions.
[0471] 5. The server uses an emotion engine to recognize emotions from the user's face and voice.
[0472] 6. Customize the display of rating results based on the emotions recognized by the emotion engine.
[0473] User terminal
[0474] 1. The user device receives the analysis results and emotion engine information sent from the server.
[0475] 2. The user device overlays the analysis results and emotion data onto frames of the video feed.
[0476] 3. The user's device will display the numerical values for technical score, composition score, and deduction points for each frame in text at the appropriate position.
[0477] 4. Dynamically change the UI based on the user's emotions, for example highlighting important information if the user is surprised.
[0478] Specific examples
[0479] For example, in a figure skating competition, the server receives a frame from the video feed of a skater performing a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points for the jump and a technical score of 7.8 points for the subsequent spin. These evaluation results are then sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice and recognizes their surprise. The user's device receives the results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." To further emphasize their surprise, the user's device also displays a comment such as "Great jump!" Spectators can view this in real time and receive interactive feedback based on the evaluation of each performance and their own emotions. This process increases the transparency of scoring criteria and further enhances the enjoyment of watching sports.
[0480] The processing flow will be explained below.
[0481] Step 1:
[0482] The server initializes a video capture device to receive the video feed from a camera or a video file. Specifically, it uses cv2.VideoCapture(0) to get real-time video data from the device.
[0483] Step 2:
[0484] The server extracts individual frames from the received video feed, using the read method of the video capture object to retrieve the image data for each frame.
[0485] Step 3:
[0486] The server passes each extracted frame to an AI evaluation module, which uses a pre-trained model to analyze the content of each frame and calculate technical and compositional scores, as well as deductions.
[0487] Step 4:
[0488] The server receives the analysis results from the AI evaluation module and converts them into a data format (e.g., JSON format), which is then prepared for transmission to the user's device.
[0489] Step 5:
[0490] The server uses an emotion engine to analyze the user's facial expressions and voice in real time to recognize their emotions. For example, it analyzes camera footage and microphone input to determine whether the user is enjoying, surprised, or dissatisfied.
[0491] Step 6:
[0492] The server customizes the displayed evaluation results based on the user's emotions. For example, if the user is surprised, it highlights specific technical or structural features.
[0493] Step 7:
[0494] The server then sends the customized evaluation results to the user's device, using a network to transfer data in real time.
[0495] Step 8:
[0496] The user device receives the analysis result data and emotion data sent from the server, and prepares an overlay display on the frame of the video feed based on this data.
[0497] Step 9:
[0498] The user's device overlays the analysis results on the frame. Specifically, the cv2.putText method is used to display the numerical information on technical points, composition points, and deductions in text format in the appropriate location.
[0499] Step 10:
[0500] The user device dynamically changes the UI based on the user's emotions. For example, if it detects that the user is surprised, it will display a comment such as "Great jump!"
[0501] Step 11:
[0502] The user device displays frames containing overlay information on the screen in real time, allowing the audience to see the evaluation results and interactive feedback in real time.
[0503] Specific examples
[0504] In a figure skating competition, the server receives a frame from the video feed of the moment an athlete performs a triple axel jump. The AI evaluation module analyzes this frame and assigns a technical score of 9.5 points for the jump. It also analyzes the frame for the spin and assigns a technical score of 7.8 points. These evaluation results are sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice to recognize surprise. The user's device receives this and overlays text such as "Triple axel: 9.5 points" and "Spin: 7.8 points" on the frame. It also displays a comment such as "Great jump!" to emphasize surprise. This allows spectators to receive interactive feedback in real time based on the evaluation of each performance and their own emotions, improving the transparency and enjoyment of sports viewing.
[0505] Example 2
[0506] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0507] Conventional sports viewing systems have the problem that it is difficult to visually display the performance evaluation of athletes in real time, and they do not provide interactive feedback according to the user's emotions. As a result, spectators cannot immediately grasp the detailed evaluation results during the performance, which limits the viewing experience.
[0508] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0509] In this invention, the server includes means for receiving a video feed, means for analyzing video data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for recognizing a user's emotions and customizing the display content, and means for overlaying the evaluation results on the frames, thereby making it possible to accurately evaluate a player's performance in real time, display the evaluation results, and provide interactive feedback according to the user's emotions.
[0510] "Video Feed" refers to a continuous stream of real-time or recorded video data.
[0511] "Motion image data" refers to digital data that contains image information that changes over time, and is typically made up of a collection of individual still images called frames.
[0512] "Technical score" is a numerical value that indicates the evaluation of the technical elements of an athlete's performance, including the accuracy and difficulty of technical movements such as jumps and spins.
[0513] "Composition score" is a numerical value that indicates an evaluation of the composition and artistic quality of a performance, and includes, for example, expressiveness and the flow of the performance.
[0514] "Deductions" are points that are subtracted from the technical or composition score due to a mistake or violation of rules.
[0515] "User emotion" refers to the subjective emotional state, such as enjoyment, surprise, or dissatisfaction, that a spectator or user feels while watching a sporting event.
[0516] "Customizing the display content" refers to dynamically changing the information and interface content displayed in response to the user's emotions.
[0517] "Overlay display" refers to a method of displaying evaluation results or additional information overlaid on the video feed.
[0518] This invention relates to a system that displays the scores of sports games in real time and customizes the display content by recognizing the emotions of users. The system is composed of a server and a user terminal.
[0519] First, the server initializes the video capture device and receives the athletes' performances in real time. Specifically, it uses OpenCV to acquire video data from IP cameras and USB cameras. The received video feed is extracted frame by frame and passed to the AI evaluation module. The AI evaluation module uses TensorFlow and PyTorch to analyze each frame and generate technical scores, composition scores, and deductions. This makes it possible to evaluate the athletes' performances in real time.
[0520] The server is also equipped with an emotion engine that recognizes the user's emotions. Using services such as Google Cloud Natural Language API and IBM Watson, it analyzes the user's facial expressions and voice to identify emotions such as whether the user is enjoying, surprised, or dissatisfied. After identifying the emotion, the server customizes the display of the evaluation results. For example, if the user is surprised, it can add a comment such as "Great jump!"
[0521] Next, the user device receives the analysis results and emotion engine information sent from the server. <canvas>Using JavaScript and JavaScript, the analysis results are overlaid on frames of the video feed. By visually displaying numerical values for technical, compositional, and deduction scores, the audience can see a detailed evaluation of each performance in real time. The UI also dynamically changes based on the user's emotions, highlighting important information as needed.
[0522] Take a figure skating competition as an example. The server receives the moment of a triple axel jump from a video feed, and the AI evaluation module assigns a technical score of 9.5 points for this jump and 7.8 points for the following spin. These evaluation results are sent to the user's device in real time, and the emotion engine recognizes that the user is surprised. The user's device receives this and displays a text overlay saying "Triple axel: 9.5 points" and "Spin: 7.8 points," along with the comment "Great jump!"
[0523] Example prompt sentence:
[0524] "Analyze a figure skater's triple axel jump in real time and rate it for technical and compositional points. Highlight results that amaze the user."
[0525] As described above, this invention allows users to check player evaluations in real time and receive interactive feedback based on their own emotions. This system will further enhance the enjoyment of watching sports.
[0526] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0527] Step 1:
[0528] The server initializes the video capture device and receives the athletes' performances in real time. Specifically, the server acquires video data using OpenCV's cv2.VideoCapture. The input is video data from the video capture device, and the output is video data divided into individual frames.
[0529] Step 2:
[0530] The server extracts frames one by one from the received video feed. Specifically, the server reads frames continuously using the OpenCV read() method and adds them to a list. The input is continuous video data from the video capture device, and the output is video data frame by frame.
[0531] Step 3:
[0532] The server passes each extracted frame to the AI evaluation module. Specifically, the server inputs the frame data into the AI evaluation module using TensorFlow or PyTorch. The input is video data for each frame, and the output is analyzed evaluation data (technical score, composition score, and deductions).
[0533] Step 4:
[0534] The AI evaluation module analyzes each frame and generates technical scores, composition scores, and deductions. Specifically, the AI model analyzes specific movements within a frame and outputs evaluation data such as "jump: 9.5 points" or "spin: 7.8 points." The input is the video data for each frame, and the output is evaluation data for each performance.
[0535] Step 5:
[0536] The server uses an emotion engine to recognize emotions from the user's face and voice. Specifically, the server analyzes the user's facial expression and voice data using Google Cloud Natural Language API and IBM Watson. The input is the user's facial expression and voice data, and the output is the user's emotional state (enjoyed, surprised, dissatisfied, etc.).
[0537] Step 6:
[0538] The display content is customized based on the emotions recognized by the emotion engine. Specifically, the server generates additional information or comments to attract the user's attention based on the analysis results. For example, if the user is surprised, it adds a comment such as "Great jump!" The input is the user's emotional state and evaluation data for each frame, and the output is customized display data.
[0539] Step 7:
[0540] The user device receives the analysis results and emotion engine information sent from the server. Specifically, the user device communicates with the server using WebSocket or RESTful API to receive data. The input is the analysis results and emotion information sent from the server, and the output is data stored on the user device.
[0541] Step 8:
[0542] The user device overlays the analysis results and emotion data on the video feed frame. <canvas>It uses elements and JavaScript to overlay analytics data on a video feed. The input is analytics results and emotion data, and the output is a visual display on the user's device.
[0543] Step 9:
[0544] The user's device displays the technical score, composition score, and deduction score values for each frame in the appropriate position. Specifically, it performs coordinate calculations to display the analysis data as text in the appropriate position within the frame. The input is the evaluation data for each frame, and the output is a text overlay display for each frame.
[0545] Step 10:
[0546] The user device dynamically changes the UI according to the user's emotions. Specifically, it uses CSS and JavaScript to change UI elements and highlight important information. For example, if the user is surprised, it changes the text color or highlights a specific comment. The input is the user's emotional data and analysis results, and the output is a dynamically changed UI display.
[0547] (Application example 2)
[0548] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0549] In conventional sports viewing systems, it was difficult for spectators to know the evaluation results of athletes' technical and compositional scores in real time, and the evaluation criteria lacked transparency. Furthermore, there was no way to provide a more interactive and engaging viewing experience by incorporating spectators' emotions and excitement. As a result, it was difficult to increase spectator satisfaction.
[0550] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0551] In this invention, the server includes means for receiving a video feed, means for analyzing video data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for overlaying the evaluation results on the frames, and means for analyzing a user's emotions and customizing the display content of the evaluation results. This allows viewers to know the evaluation results, such as technical scores and composition scores, in real time when watching a sports game, and further provides an interactive viewing experience that responds to the viewers' emotions.
[0552] A "video feed" is a continuous video signal coming from a camera or other video capture device.
[0553] A "frame" refers to an individual still image within a video feed, with a series of frames forming a moving image.
[0554] "Video data" refers to the digital video data that comprises each frame in a video feed.
[0555] "Technical score" is a numerical evaluation of an athlete's technical performance in a sporting event.
[0556] "Composition points" are a numerical evaluation of the composition and expressiveness of a performance in a sporting event.
[0557] "Deduction points" refers to the evaluation of a sporting event in which negative points are given for technical errors or violations of rules.
[0558] "Evaluation results" refers to the overall score including technical points, composition points, and deductions.
[0559] "Overlay display" refers to a technique for displaying additional information overlaid on top of a video frame.
[0560] "User's emotion" refers to the subjective emotional state that the user feels while watching a sporting event, and is analyzed from facial expressions and voice.
[0561] "Emotion engine" refers to a software engine that recognizes and analyzes emotions from the user's face and voice.
[0562] "UI" stands for user interface and refers to the screen layout and design that allows users to interact with systems and applications.
[0563] "Real-time" refers to information and data processing occurring in real time, with results reflected almost immediately.
[0564] A system for implementing this invention consists of a server and a user terminal. The server receives a video feed and analyzes the video data frame by frame to generate evaluation results in terms of technical points, composition points, and deductions. It also analyzes the user's emotions and customizes the display of the evaluation results. The user terminal displays the evaluation results as an overlay, visually presenting them to the user.
[0565] Specifically, the server receives live feeds from cameras and video capture devices. This live feed contains the athletes' performances and analyzes them frame by frame. An AI evaluation module using TensorFlow is used for the analysis, and evaluation results are generated in real time, including technical scores, compositional scores, and deductions. These evaluation results are then sent to the user's device.
[0566] The server is equipped with an emotion engine that recognizes emotions from viewers' faces and voices. This emotion engine uses VaderSentiment and TensorFlow emotion recognition models. Viewers' facial expressions and voice data are analyzed using the dlib library to grasp the user's emotional state in real time. As a result, the display of the evaluation results is customized, such as highlighting important information if the user is surprised.
[0567] The user's device receives the analysis results and emotion data sent from the server. A Python-based GUI application is installed on the user's device, which overlays the live feed frames with information on technical and compositional points, as well as deductions. Furthermore, the application dynamically displays comments and supplementary information based on the user's emotions.
[0568] As a concrete example, in a figure skating competition, the server analyzes the frame at the moment when the skater performs a triple axel jump. The AI evaluation module generates a technical score of 9.5 points for this jump, and a technical score of 7.8 points for the subsequent spin. These evaluation results are sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice and recognizes that the user is surprised. The user's device receives this and overlays the frame with text such as "Triple axel: 9.5 points" and "Spin: 7.8 points," along with additional comments such as "Great jump!"
[0569] An example prompt is:
[0570] "Design a system that displays technical and compositional scores in real time based on viewer emotions during live-streamed sporting events. Utilize an AI evaluation module and emotion recognition engine to dynamically display cheering comments, such as 'Great jump!', when a viewer expresses surprise."
[0571] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0572] Step 1:
[0573] The server receives the video feed from the camera or video capture device. Here, the input is a live video feed, and each frame of the video feed is continuously sent to the server. This step captures the athlete's performance in real time. The output is a continuous video signal.
[0574] Step 2:
[0575] The server extracts each frame from the received video feed. The input is the video feed, and data is extracted frame by frame. Here, the actual processing is performed to extract the video data obtained from the video capture device one frame at a time. The output is frame data, which are individual still images.
[0576] Step 3:
[0577] The server passes the extracted frame data to the AI evaluation module. The input is the frame data, which is analyzed frame by frame using TensorFlow. The module calculates the technical score, composition score, and deduction points, and outputs the evaluation results as numerical data.
[0578] Step 4:
[0579] The server is equipped with an emotion engine that analyzes emotions from the user's facial expressions and voice. The input is the user's facial expression data and voice data, which are analyzed using the dlib library and VaderSentiment. In this step, specific processing is performed to identify the user's emotional state, such as whether they are surprised, amused, or frustrated. The output is the user's emotional state data.
[0580] Step 5:
[0581] The server combines the evaluation results and the user's emotional state to customize the content of the evaluation results. The input is the numerical evaluation results and emotional state data, and the display content is dynamically changed based on these. In this step, specific operations are performed, such as adding a comment such as "Great jump!" if the user is surprised. The output is customized evaluation result data.
[0582] Step 6:
[0583] The server sends the customized evaluation result data to the user terminal. The input is the customized evaluation result data, which is sent to the user terminal via the network. In this step, the specific operation of data transfer is performed. The output is the evaluation result data sent to the user terminal.
[0584] Step 7:
[0585] The user device uses the received evaluation result data and emotion data to overlay a frame of the video feed. The input is the customized evaluation result data and emotion data, and specific processing is performed to display text and comments based on these. The output is a visual display result.
[0586] Step 8:
[0587] The user device dynamically changes the UI according to the user's emotional response. The input is real-time emotional data, and specific actions are taken to change the interface and presented information based on this. The output is an updated UI.
[0588] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0589] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0590] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0591] [Third embodiment]
[0592] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0593] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0594] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0595] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0596] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0597] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0598] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0599] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0600] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0601] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0602] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0603] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0604] The present invention relates to a system for displaying real-time scoring results during sports viewing. The system receives a video feed, analyzes it, generates technical scores, composition scores, and deductions, and visually displays the results. Specific embodiments of the system are described below.
[0605] System Overview
[0606] server
[0607] First, the server receives a video feed from a camera or a video file. The video feed contains footage of the athletes' performances and is processed frame by frame. The server initializes the video capture device and continuously receives the video data.
[0608] The server then extracts each frame from the received video feed and passes it to an AI evaluation module, which analyzes each frame and generates a technical score, composition score, and deductions for mistakes. This analysis result is obtained for each frame, allowing for real-time evaluation.
[0609] User terminal
[0610] The user's device receives the analysis results sent from the server and displays them as an overlay on the video feed frame. The numerical values for technical scores, composition scores, and deductions are displayed as text in the appropriate position within the frame. This allows the audience to visually check the evaluation score for each performance in real time, and intuitively understand the scoring criteria.
[0611] Program processing
[0612] server
[0613] 1. The server continues to receive the video feed from the video capture device.
[0614] 2. The server extracts each frame from the received video feed.
[0615] 3. The server passes each extracted frame to the AI evaluation module.
[0616] 4. The AI evaluation module analyzes each frame and generates an evaluation result including technical points, composition points, and deductions.
[0617] User terminal
[0618] 1. The user device receives the analysis results sent from the server.
[0619] 2. The user device overlays the analysis results on frames of the video feed.
[0620] 3. The user's device will display the numerical values for technical score, composition score, and deduction points for each frame in text at the appropriate position.
[0621] Specific examples
[0622] For example, in a figure skating competition, the server receives a frame from the video feed in which an athlete performs a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points to this jump, and a technical score of 7.8 points to the subsequent spin. The user device receives these evaluation results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can see this in real time and understand the performance evaluation on the spot. This improves the transparency and enjoyment of watching sports.
[0623] The processing flow will be explained below.
[0624] Step 1:
[0625] The server initializes the video capture device to receive the video feed from the camera or video file. A video capture object is created using cv2.VideoCapture(0), which starts capturing real-time video data.
[0626] Step 2:
[0627] The server reads the video data frame by frame from the received video feed, specifically by using the read method of the video capture object to extract the frame-by-frame images.
[0628] Step 3:
[0629] The server passes the acquired frames to the AI evaluation module, which receives the images frame by frame as preprocessing and extracts features related to the player's movements from the input.
[0630] Step 4:
[0631] The server receives the analysis results from the AI evaluation module, which calculates the technical score, composition score, and deduction score for each frame and returns the evaluation results including this data.
[0632] Step 5:
[0633] The server formats the evaluation results and converts them into a data format for sending to the user's device. Specifically, the evaluation results are converted into a flexible data format such as JSON.
[0634] Step 6:
[0635] The user terminal receives the analysis result data sent from the server, analyzes this data, and extracts the necessary information.
[0636] Step 7:
[0637] The user device overlays the analysis results on the video feed frame, using the cv2.putText method to display numerical information such as technical scores, composition scores, and deductions in the appropriate positions on the frame.
[0638] Step 8:
[0639] The user device displays the frame containing the overlay information on the screen in real time using the cv2.imshow method, allowing the audience to instantly see the evaluation results on the screen.
[0640] Specific examples
[0641] In a figure skating competition, the server receives the frame at the moment a triple axel jump is performed. The AI evaluation module analyzes this frame and assigns a technical score of 9.5 points for the jump. It also analyzes the frame for the subsequent spin and assigns a technical score of 7.8 points. These evaluation results are sent from the server to the user's device, which receives them. The user's device displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can view this in real time and instantly understand the evaluation of each performance. This process improves the transparency of scoring criteria and deepens understanding of sports spectators.
[0642] Example 1
[0643] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0644] Conventional sports viewing systems make it difficult for spectators to visually understand the technical and compositional scores of athletes in real time. Furthermore, the lack of transparency and speed in the evaluation process reduces the enjoyment of watching the game.
[0645] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0646] In this invention, the server includes means for receiving a video feed from a video capture device, means for extracting frames from the video feed and analyzing the video data to generate evaluation results regarding technical scores, composition scores, and deductions, means for overlaying the evaluation results on the frames in real time, and means for visually displaying the evaluation results in text format, thereby enabling spectators to visually check the evaluation scores in real time and enjoy fast and transparent evaluation.
[0647] "Video capture device" means equipment or software for receiving a video feed, including cameras and video streaming devices.
[0648] A "video feed" is a continuously distributed stream of moving image data, including footage of sporting matches and events.
[0649] "Frames" are the individual still images that make up a video feed and are played back in succession to form a moving image.
[0650] "Motion image data" is continuous image data consisting of multiple frames, and is used to analyze motion and movement.
[0651] "Technical score" is a technical evaluation score of an athlete's performance, based on the degree of perfection and difficulty of a particular movement.
[0652] "Composition score" is an evaluation score based on the composition and direction of the athlete's entire performance, taking into account creativity and expressiveness.
[0653] "Deductions" are points deducted from an athlete's technical and compositional scores for mistakes or incorrect movements during their performance.
[0654] The "evaluation result" is an analysis result including technical points, composition points, and deductions, and is generated for each frame.
[0655] "Overlay display" refers to the process of displaying the evaluation results overlaid on frames of the video feed, allowing the evaluation to be visually confirmed in real time.
[0656] "Text format" is a format in which information is displayed as text, and presents numerical values, evaluation points, etc. in a visually easy-to-understand manner.
[0657] A "deep learning evaluation module" is a module that uses artificial intelligence technology to analyze video data and evaluate technical and compositional aspects, and includes neural networks.
[0658] The present invention relates to a system for displaying real-time scoring results during sports viewing. The system receives video feeds from a video capture device, analyzes the video feeds, evaluates technical scores, composition scores, and deductions, and visually displays the evaluation results. Specific embodiments of the system are described below.
[0659] server
[0660] First, the server initializes a video capture device (e.g., HD camera, IP camera) as a video acquisition device and receives the video feed. The video feed contains footage of the athletes' performances, which is processed frame by frame. The server extracts the frames using an image processing library such as OpenCV.
[0661] Next, the server inputs each extracted frame into a deep learning evaluation module (e.g., a model using TensorFlow or PyTorch). This deep learning evaluation module analyzes each frame and evaluates technical points, composition points, and deductions. Analysis results are generated for each frame, allowing for real-time evaluation.
[0662] Specifically, the server receives a frame of a figure skating competition in which a skater performs a triple axel jump. This frame is preprocessed and sent to the evaluation module for technical evaluation. The evaluation module assigns a technical score of 9.5 points to the jump and 7.8 points to the subsequent spin.
[0663] User terminal
[0664] The user device receives the analysis results sent from the server. This data is sent and received using a real-time communication protocol such as WebSocket. Based on the received analysis results, the user device displays the evaluation results on the video feed frames. Specifically, numerical information such as technical scores, composition scores, and deductions is overlaid on the video frames using the HTML5 Canvas element and OpenGL.
[0665] Furthermore, the user's device will display the evaluation score in text format at the appropriate position for each frame. For example, text such as "Triple Axel: 9.5 points" or "Spin: 7.8 points" will be displayed in the bottom left or bottom right. This allows the audience to visually check the evaluation score in real time.
[0666] This system will improve the transparency of sports viewing and increase the enjoyment of watching, and will also enable spectators to intuitively understand the evaluation of performances in real time.
[0667] Prompt Sentence Examples
[0668] Below are some example prompts to input to the generative AI model:
[0669] "Describe an AI system that analyzes a video feed of figure skating performances and displays the technical and compositional scores for each performance in real time."
[0670] "Please tell us the specific technical details of a sports viewing system that displays the types of moves performed by athletes and their evaluation scores in real time."
[0671] The above is a specific embodiment for carrying out the present invention.
[0672] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0673] Step 1:
[0674] The server receives a video feed from a video capture device. The video capture device can be an HD camera or an IP camera. The input is the video feed from the video capture device, and the output is video image data divided into frames. This video feed is continuously captured and stored in the server's memory.
[0675] Step 2:
[0676] The server extracts individual frames from the video feed. It uses an image processing library such as OpenCV to acquire each frame. The input is the video feed received in step 1, and the output is image data for each frame. Specifically, it uses the OpenCV read function to extract the frames.
[0677] Step 3:
[0678] The server passes each extracted frame to a deep learning evaluation module. Using a deep learning framework such as TensorFlow or PyTorch, the module extracts and evaluates the frame's features. The input is the frame image extracted in step 2, and the output is the evaluation results, including technical scores, composition scores, and deductions. A convolutional neural network (CNN) is used for this evaluation to analyze the movement.
[0679] Step 4:
[0680] The deep learning evaluation module analyzes each frame and generates evaluation results in terms of technical points, composition points, and deductions. Specifically, it uses CNN to identify actions within the frame and calculate a score for each action. The input here is the frame image passed in step 3, and the output is the evaluation result for each frame.
[0681] Step 5:
[0682] The server sends the generated evaluation results to the user's device in real time using WebSocket. The input is the evaluation results generated in step 4, and the output is the analysis data sent to the user's device.
[0683] Step 6:
[0684] The user terminal receives the analysis results sent from the server. The input here is the evaluation results sent in step 5, and the output is the evaluation data stored in the memory on the user terminal.
[0685] Step 7:
[0686] The user device overlays the received analysis results on the video feed frame. Using the HTML5 Canvas element or OpenGL, the numerical information of the evaluation results is displayed overlaid on the video frame. The input is the evaluation results received in step 6, and the output is the overlaid video frame. Specific operations include displaying text in the appropriate position.
[0687] Step 8:
[0688] The user device displays the technical score, composition score, and deduction scores in text format at the appropriate position on the video frame. The input is the frame overlaid in step 7, and the output is real-time video that allows the audience to visually check the evaluation scores. For example, text information such as "Triple Axel: 9.5 points" or "Spin: 7.8 points" is displayed.
[0689] This series of processes allows spectators to visually check the evaluation scores for athletes' performances in real time, improving the transparency and enjoyment of watching sports.
[0690] (Application example 1)
[0691] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0692] With conventional sports viewing systems, it was difficult for spectators to check the evaluation of performances and plays in real time. In particular, there were limited means for them to visually understand the evaluation of technical points, composition points, and deductions. This reduced the transparency and enjoyment of the viewing experience, and prevented spectators from fully experiencing the appeal of sports. Another issue was the lack of systems that took into account use on mobile devices such as smartphones.
[0693] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0694] In this invention, the server includes means for receiving a video feed, means for analyzing moving image data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for overlaying and displaying the evaluation results on the frames, means for receiving a video feed from a live streaming server, means for analyzing the video frames with an AI evaluation module and generating evaluation results, and means for overlaying and displaying the analysis results on the video feed in real time, thereby enabling spectators to visually understand the evaluation results of sports performances and plays being live-streamed in real time using their smartphones.
[0695] A "video feed" is a continuous stream of image data transmitted from a camera or other video capture device.
[0696] A "frame" is an individual still image in a video feed that is displayed in succession to form a moving image.
[0697] "Motion image data" refers to video data consisting of a series of frames obtained from a video feed.
[0698] "Technical points" are the points awarded to athletes in sports competitions based on their technical performance and skills.
[0699] "Composition points" are points given in sports competitions to evaluate the overall composition and artistry of an athlete's performance.
[0700] "Deductions" are points that are deducted in sports competitions when a player commits a mistake or acts in violation of the rules.
[0701] "AI Evaluation Module" means a software module that uses artificial intelligence technology to analyze video frames and generate technical scores, composition scores, and deduction scores.
[0702] A "live streaming server" is a server for delivering video feeds in real time over the Internet.
[0703] "Overlay display" is a technique for displaying additional information superimposed on a video frame.
[0704] "Real-time" means processed and displayed immediately, without delay or lag.
[0705] A "smartphone" is a portable information terminal that has mobile phone functions and can be used by installing additional applications.
[0706] "Analysis results" are the evaluation data of technical points, composition points, and deduction points generated by the AI evaluation module.
[0707] "Visual presentation" means displaying information in a way that is easy for the audience to see, and is a technique that allows information to be intuitively understood.
[0708] This invention relates to a system for displaying evaluation results in real time while watching sports. The system can receive a video feed, analyze it, generate evaluations of technical points, composition points, and deductions, and display the results visually.
[0709] System Overview
[0710] server
[0711] The server receives the video feed from the live streaming server and extracts video data from the video feed frame by frame. It initializes the video capture device and continuously receives video data. The server then extracts each frame from the received video feed and passes it to the AI evaluation module. The AI evaluation module analyzes each frame and generates technical scores, composition scores, and deductions for mistakes. The analysis results are obtained for each frame, allowing for real-time evaluation.
[0712] User terminal
[0713] The user's device receives the analysis results sent from the server and displays them as an overlay on the video feed frame. The numerical values for technical points, composition points, and deductions are displayed as text in the appropriate position within the frame. This allows spectators to visually check the evaluation scores for each performance and play in real time, and to intuitively understand the evaluation criteria.
[0714] Program processing
[0715] The implementation of this system requires several key software and hardware components: OpenCV is used to sample the video feed and extract frames, TensorFlow is used to generate scores for the AI evaluation module, and the phone's native UI framework (e.g., UIKit for iOS, Jetpack Compose for Android) is used to render the overlay display.
[0716] Specific hardware and software
[0717] Smartphone: iOS or Android device
[0718] Live Streaming Server: Streams video feeds using ffmpeg
[0719] AI evaluation module: Machine learning libraries such as TensorFlow
[0720] Frame analysis software: OpenCV
[0721] UI Framework: UIKit or Jetpack Compose
[0722] Specific examples
[0723] For example, in a figure skating competition, the server receives a frame from the video feed of the moment an athlete performs a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points for this jump and a technical score of 7.8 points for the subsequent spin. The user's device receives these evaluation results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can see this in real time and understand the performance evaluation on the spot. This improves the transparency and enjoyment of watching sports.
[0724] Prompt Sentence Examples
[0725] Create an application that displays real-time scoring results while watching sports. The app receives a video feed from a live streaming server and analyzes each frame with a TensorFlow model to generate and display technical scores, composition scores, and deductions. The software used is ffmpeg, OpenCV, and TensorFlow.
[0726] By implementing this invention, spectators can use their smartphones to visually understand the evaluation results of live-streamed sports performances and plays in real time.
[0727] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0728] Step 1:
[0729] The server receives a video feed from the live streaming server. It initializes a video capture device and continuously receives video data. The input is the video feed from the live streaming server, and the output is a stream of video frames. Specifically, the server initializes a video capture device and continuously processes the received video stream.
[0730] Step 2:
[0731] The server extracts each frame from the received video feed. It uses OpenCV to extract the frames and convert them into individual still image data. The input is the video feed, and the output is still image data for each frame. Specifically, it uses OpenCV functions to extract still images for each frame from the video stream.
[0732] Step 3:
[0733] The server passes each extracted frame to an AI evaluation module, which uses TensorFlow to analyze the frames and generate evaluations for technical, compositional, and deduction points. The input is still image data for each frame, and the output is the evaluation results for technical, compositional, and deduction points. Specifically, a pre-trained model is used to analyze the data for each frame and calculate the points.
[0734] Step 4:
[0735] The server transmits the evaluation results generated by the AI evaluation module to the user terminal. The evaluation results are transmitted in real time. The input is the evaluation result for each frame, and the output is the evaluation data transmitted over the network. Specifically, the server converts the evaluation results into an appropriate format and transmits them over the network.
[0736] Step 5:
[0737] The user device overlays the analysis results received from the server on the video feed frame. The input is the evaluation results sent from the server, and the output is the overlaid video frame. Specifically, the device's native UI framework is used to display the evaluation results in text on the frame in real time.
[0738] Step 6:
[0739] The user checks the evaluation results in real time using a smartphone. The input is the video feed displayed on the user's device, and the output is the user's visual understanding of the information. Specifically, the user checks the performances and plays in the video stream along with the evaluation scores overlaid in real time.
[0740] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0741] The present invention relates to a system that displays real-time scoring results during sports viewing and customizes the display content by recognizing the user's emotions. The system receives a video feed, analyzes it, generates technical scores, composition scores, and deductions, and visually displays the results. It also incorporates an emotion engine that recognizes the user's emotions. Specific embodiments of the system are described below.
[0742] System Overview
[0743] server
[0744] First, the server receives a video feed from a camera or a video file. The video feed contains footage of the athletes' performances and is processed frame by frame. The server initializes the video capture device and continuously receives the video data.
[0745] The server then extracts each frame from the received video feed and passes it to an AI evaluation module, which analyzes each frame and generates a technical score, composition score, and deductions for mistakes. This analysis result is obtained for each frame, allowing for real-time evaluation.
[0746] Furthermore, the server is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's facial expressions and voice to identify emotions such as enjoyment, surprise, or dissatisfaction. Based on the recognized emotion, the server customizes the display of the evaluation results. For example, if the user shows interest, it can provide additional detailed information.
[0747] User terminal
[0748] The user's device receives the analysis results sent from the server as well as information from the emotion engine. Based on this, an overlay display is performed on the video feed frame. As text, numerical values for technical points, composition points, and deductions are displayed in appropriate positions within the frame. Furthermore, a UI (user interface) can be constructed according to the user's emotions, and the displayed content can be dynamically changed.
[0749] Program processing
[0750] server
[0751] 1. The server continues to receive the video feed from the video capture device.
[0752] 2. The server extracts each frame from the received video feed.
[0753] 3. The server passes each extracted frame to the AI evaluation module.
[0754] 4. The AI evaluation module analyzes each frame and generates an evaluation result including technical points, composition points, and deductions.
[0755] 5. The server uses an emotion engine to recognize emotions from the user's face and voice.
[0756] 6. Customize the display of rating results based on the emotions recognized by the emotion engine.
[0757] User terminal
[0758] 1. The user device receives the analysis results and emotion engine information sent from the server.
[0759] 2. The user device overlays the analysis results and emotion data onto frames of the video feed.
[0760] 3. The user's device will display the numerical values for technical score, composition score, and deduction points for each frame in text at the appropriate position.
[0761] 4. Dynamically change the UI based on the user's emotions, for example highlighting important information if the user is surprised.
[0762] Specific examples
[0763] For example, in a figure skating competition, the server receives a frame from the video feed of a skater performing a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points for the jump and a technical score of 7.8 points for the subsequent spin. These evaluation results are then sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice and recognizes their surprise. The user's device receives the results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." To further emphasize their surprise, the user's device also displays a comment such as "Great jump!" Spectators can view this in real time and receive interactive feedback based on the evaluation of each performance and their own emotions. This process increases the transparency of scoring criteria and further enhances the enjoyment of watching sports.
[0764] The processing flow will be explained below.
[0765] Step 1:
[0766] The server initializes a video capture device to receive the video feed from a camera or a video file. Specifically, it uses cv2.VideoCapture(0) to get real-time video data from the device.
[0767] Step 2:
[0768] The server extracts individual frames from the received video feed, using the read method of the video capture object to retrieve the image data for each frame.
[0769] Step 3:
[0770] The server passes each extracted frame to an AI evaluation module, which uses a pre-trained model to analyze the content of each frame and calculate technical and compositional scores, as well as deductions.
[0771] Step 4:
[0772] The server receives the analysis results from the AI evaluation module and converts them into a data format (e.g., JSON format), which is then prepared for transmission to the user's device.
[0773] Step 5:
[0774] The server uses an emotion engine to analyze the user's facial expressions and voice in real time to recognize their emotions. For example, it analyzes camera footage and microphone input to determine whether the user is enjoying, surprised, or dissatisfied.
[0775] Step 6:
[0776] The server customizes the displayed evaluation results based on the user's emotions. For example, if the user is surprised, it highlights specific technical or structural features.
[0777] Step 7:
[0778] The server then sends the customized evaluation results to the user's device, using a network to transfer data in real time.
[0779] Step 8:
[0780] The user device receives the analysis result data and emotion data sent from the server, and prepares an overlay display on the frame of the video feed based on this data.
[0781] Step 9:
[0782] The user's device overlays the analysis results on the frame. Specifically, the cv2.putText method is used to display the numerical information on technical points, composition points, and deductions in text format in the appropriate location.
[0783] Step 10:
[0784] The user device dynamically changes the UI based on the user's emotions. For example, if it detects that the user is surprised, it will display a comment such as "Great jump!"
[0785] Step 11:
[0786] The user device displays frames containing overlay information on the screen in real time, allowing the audience to see the evaluation results and interactive feedback in real time.
[0787] Specific examples
[0788] In a figure skating competition, the server receives a frame from the video feed of the moment an athlete performs a triple axel jump. The AI evaluation module analyzes this frame and assigns a technical score of 9.5 points for the jump. It also analyzes the frame for the spin and assigns a technical score of 7.8 points. These evaluation results are sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice to recognize surprise. The user's device receives this and overlays text such as "Triple axel: 9.5 points" and "Spin: 7.8 points" on the frame. It also displays a comment such as "Great jump!" to emphasize surprise. This allows spectators to receive interactive feedback in real time based on the evaluation of each performance and their own emotions, improving the transparency and enjoyment of sports viewing.
[0789] Example 2
[0790] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0791] Conventional sports viewing systems have the problem that it is difficult to visually display the performance evaluation of athletes in real time, and they do not provide interactive feedback according to the user's emotions. As a result, spectators cannot immediately grasp the detailed evaluation results during the performance, which limits the viewing experience.
[0792] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0793] In this invention, the server includes means for receiving a video feed, means for analyzing video data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for recognizing a user's emotions and customizing the display content, and means for overlaying the evaluation results on the frames, thereby making it possible to accurately evaluate a player's performance in real time, display the evaluation results, and provide interactive feedback according to the user's emotions.
[0794] "Video Feed" refers to a continuous stream of real-time or recorded video data.
[0795] "Motion image data" refers to digital data that contains image information that changes over time, and is typically made up of a collection of individual still images called frames.
[0796] "Technical score" is a numerical value that indicates the evaluation of the technical elements of an athlete's performance, including the accuracy and difficulty of technical movements such as jumps and spins.
[0797] "Composition score" is a numerical value that indicates an evaluation of the composition and artistic quality of a performance, and includes, for example, expressiveness and the flow of the performance.
[0798] "Deductions" are points that are subtracted from the technical or composition score due to a mistake or violation of rules.
[0799] "User emotion" refers to the subjective emotional state, such as enjoyment, surprise, or dissatisfaction, that a spectator or user feels while watching a sporting event.
[0800] "Customizing the display content" refers to dynamically changing the information and interface content displayed in response to the user's emotions.
[0801] "Overlay display" refers to a method of displaying evaluation results or additional information overlaid on the video feed.
[0802] This invention relates to a system that displays the scores of sports games in real time and customizes the display content by recognizing the emotions of users. The system is composed of a server and a user terminal.
[0803] First, the server initializes the video capture device and receives the athletes' performances in real time. Specifically, it uses OpenCV to acquire video data from IP cameras and USB cameras. The received video feed is extracted frame by frame and passed to the AI evaluation module. The AI evaluation module uses TensorFlow and PyTorch to analyze each frame and generate technical scores, composition scores, and deductions. This makes it possible to evaluate the athletes' performances in real time.
[0804] The server is also equipped with an emotion engine that recognizes the user's emotions. Using services such as Google Cloud Natural Language API and IBM Watson, it analyzes the user's facial expressions and voice to identify emotions such as whether the user is enjoying, surprised, or dissatisfied. After identifying the emotion, the server customizes the display of the evaluation results. For example, if the user is surprised, it can add a comment such as "Great jump!"
[0805] Next, the user device receives the analysis results and emotion engine information sent from the server. <canvas>Using JavaScript and JavaScript, the analysis results are overlaid on frames of the video feed. By visually displaying numerical values for technical, compositional, and deduction scores, the audience can see a detailed evaluation of each performance in real time. The UI also dynamically changes based on the user's emotions, highlighting important information as needed.
[0806] Take a figure skating competition as an example. The server receives the moment of a triple axel jump from a video feed, and the AI evaluation module assigns a technical score of 9.5 points for this jump and 7.8 points for the following spin. These evaluation results are sent to the user's device in real time, and the emotion engine recognizes that the user is surprised. The user's device receives this and displays a text overlay saying "Triple axel: 9.5 points" and "Spin: 7.8 points," along with the comment "Great jump!"
[0807] Example prompt sentence:
[0808] "Analyze a figure skater's triple axel jump in real time and rate it for technical and compositional points. Highlight results that amaze the user."
[0809] As described above, this invention allows users to check player evaluations in real time and receive interactive feedback based on their own emotions. This system will further enhance the enjoyment of watching sports.
[0810] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0811] Step 1:
[0812] The server initializes the video capture device and receives the athletes' performances in real time. Specifically, the server acquires video data using OpenCV's cv2.VideoCapture. The input is video data from the video capture device, and the output is video data divided into individual frames.
[0813] Step 2:
[0814] The server extracts frames one by one from the received video feed. Specifically, the server reads frames continuously using the OpenCV read() method and adds them to a list. The input is continuous video data from the video capture device, and the output is video data frame by frame.
[0815] Step 3:
[0816] The server passes each extracted frame to the AI evaluation module. Specifically, the server inputs the frame data into the AI evaluation module using TensorFlow or PyTorch. The input is video data for each frame, and the output is analyzed evaluation data (technical score, composition score, and deductions).
[0817] Step 4:
[0818] The AI evaluation module analyzes each frame and generates technical scores, composition scores, and deductions. Specifically, the AI model analyzes specific movements within a frame and outputs evaluation data such as "jump: 9.5 points" or "spin: 7.8 points." The input is the video data for each frame, and the output is evaluation data for each performance.
[0819] Step 5:
[0820] The server uses an emotion engine to recognize emotions from the user's face and voice. Specifically, the server analyzes the user's facial expression and voice data using Google Cloud Natural Language API and IBM Watson. The input is the user's facial expression and voice data, and the output is the user's emotional state (enjoyed, surprised, dissatisfied, etc.).
[0821] Step 6:
[0822] The display content is customized based on the emotions recognized by the emotion engine. Specifically, the server generates additional information or comments to attract the user's attention based on the analysis results. For example, if the user is surprised, it adds a comment such as "Great jump!" The input is the user's emotional state and evaluation data for each frame, and the output is customized display data.
[0823] Step 7:
[0824] The user device receives the analysis results and emotion engine information sent from the server. Specifically, the user device communicates with the server using WebSocket or RESTful API to receive data. The input is the analysis results and emotion information sent from the server, and the output is data stored on the user device.
[0825] Step 8:
[0826] The user device overlays the analysis results and emotion data on the video feed frame. <canvas>It uses elements and JavaScript to overlay analytics data on a video feed. The input is analytics results and emotion data, and the output is a visual display on the user's device.
[0827] Step 9:
[0828] The user's device displays the technical score, composition score, and deduction score values for each frame in the appropriate position. Specifically, it performs coordinate calculations to display the analysis data as text in the appropriate position within the frame. The input is the evaluation data for each frame, and the output is a text overlay display for each frame.
[0829] Step 10:
[0830] The user device dynamically changes the UI according to the user's emotions. Specifically, it uses CSS and JavaScript to change UI elements and highlight important information. For example, if the user is surprised, it changes the text color or highlights a specific comment. The input is the user's emotional data and analysis results, and the output is a dynamically changed UI display.
[0831] (Application example 2)
[0832] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0833] In conventional sports viewing systems, it was difficult for spectators to know the evaluation results of athletes' technical and compositional scores in real time, and the evaluation criteria lacked transparency. Furthermore, there was no way to provide a more interactive and engaging viewing experience by incorporating spectators' emotions and excitement. As a result, it was difficult to increase spectator satisfaction.
[0834] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0835] In this invention, the server includes means for receiving a video feed, means for analyzing video data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for overlaying the evaluation results on the frames, and means for analyzing a user's emotions and customizing the display content of the evaluation results. This allows viewers to know the evaluation results, such as technical scores and composition scores, in real time when watching a sports game, and further provides an interactive viewing experience that responds to the viewers' emotions.
[0836] A "video feed" is a continuous video signal coming from a camera or other video capture device.
[0837] A "frame" refers to an individual still image within a video feed, with a series of frames forming a moving image.
[0838] "Video data" refers to the digital video data that comprises each frame in a video feed.
[0839] "Technical score" is a numerical evaluation of an athlete's technical performance in a sporting event.
[0840] "Composition points" are a numerical evaluation of the composition and expressiveness of a performance in a sporting event.
[0841] "Deduction points" refers to the evaluation of a sporting event in which negative points are given for technical errors or violations of rules.
[0842] "Evaluation results" refers to the overall score including technical points, composition points, and deductions.
[0843] "Overlay display" refers to a technique for displaying additional information overlaid on top of a video frame.
[0844] "User's emotion" refers to the subjective emotional state that the user feels while watching a sporting event, and is analyzed from facial expressions and voice.
[0845] "Emotion engine" refers to a software engine that recognizes and analyzes emotions from the user's face and voice.
[0846] "UI" stands for user interface and refers to the screen layout and design that allows users to interact with systems and applications.
[0847] "Real-time" refers to information and data processing occurring in real time, with results reflected almost immediately.
[0848] A system for implementing this invention consists of a server and a user terminal. The server receives a video feed and analyzes the video data frame by frame to generate evaluation results in terms of technical points, composition points, and deductions. It also analyzes the user's emotions and customizes the display of the evaluation results. The user terminal displays the evaluation results as an overlay, visually presenting them to the user.
[0849] Specifically, the server receives live feeds from cameras and video capture devices. This live feed contains the athletes' performances and analyzes them frame by frame. An AI evaluation module using TensorFlow is used for the analysis, and evaluation results are generated in real time, including technical scores, compositional scores, and deductions. These evaluation results are then sent to the user's device.
[0850] The server is equipped with an emotion engine that recognizes emotions from viewers' faces and voices. This emotion engine uses VaderSentiment and TensorFlow emotion recognition models. Viewers' facial expressions and voice data are analyzed using the dlib library to grasp the user's emotional state in real time. As a result, the display of the evaluation results is customized, such as highlighting important information if the user is surprised.
[0851] The user's device receives the analysis results and emotion data sent from the server. A Python-based GUI application is installed on the user's device, which overlays the live feed frames with information on technical and compositional points, as well as deductions. Furthermore, the application dynamically displays comments and supplementary information based on the user's emotions.
[0852] As a concrete example, in a figure skating competition, the server analyzes the frame at the moment when the skater performs a triple axel jump. The AI evaluation module generates a technical score of 9.5 points for this jump, and a technical score of 7.8 points for the subsequent spin. These evaluation results are sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice and recognizes that the user is surprised. The user's device receives this and overlays the frame with text such as "Triple axel: 9.5 points" and "Spin: 7.8 points," along with additional comments such as "Great jump!"
[0853] An example prompt is:
[0854] "Design a system that displays technical and compositional scores in real time based on viewer emotions during live-streamed sporting events. Utilize an AI evaluation module and emotion recognition engine to dynamically display cheering comments, such as 'Great jump!', when a viewer expresses surprise."
[0855] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0856] Step 1:
[0857] The server receives the video feed from the camera or video capture device. Here, the input is a live video feed, and each frame of the video feed is continuously sent to the server. This step captures the athlete's performance in real time. The output is a continuous video signal.
[0858] Step 2:
[0859] The server extracts each frame from the received video feed. The input is the video feed, and data is extracted frame by frame. Here, the actual processing is performed to extract the video data obtained from the video capture device one frame at a time. The output is frame data, which are individual still images.
[0860] Step 3:
[0861] The server passes the extracted frame data to the AI evaluation module. The input is the frame data, which is analyzed frame by frame using TensorFlow. The module calculates the technical score, composition score, and deduction points, and outputs the evaluation results as numerical data.
[0862] Step 4:
[0863] The server is equipped with an emotion engine that analyzes emotions from the user's facial expressions and voice. The input is the user's facial expression data and voice data, which are analyzed using the dlib library and VaderSentiment. In this step, specific processing is performed to identify the user's emotional state, such as whether they are surprised, amused, or frustrated. The output is the user's emotional state data.
[0864] Step 5:
[0865] The server combines the evaluation results and the user's emotional state to customize the content of the evaluation results. The input is the numerical evaluation results and emotional state data, and the display content is dynamically changed based on these. In this step, specific operations are performed, such as adding a comment such as "Great jump!" if the user is surprised. The output is customized evaluation result data.
[0866] Step 6:
[0867] The server sends the customized evaluation result data to the user terminal. The input is the customized evaluation result data, which is sent to the user terminal via the network. In this step, the specific operation of data transfer is performed. The output is the evaluation result data sent to the user terminal.
[0868] Step 7:
[0869] The user device uses the received evaluation result data and emotion data to overlay a frame of the video feed. The input is the customized evaluation result data and emotion data, and specific processing is performed to display text and comments based on these. The output is a visual display result.
[0870] Step 8:
[0871] The user device dynamically changes the UI according to the user's emotional response. The input is real-time emotional data, and specific actions are taken to change the interface and presented information based on this. The output is an updated UI.
[0872] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0873] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0874] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0875] [Fourth embodiment]
[0876] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0877] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0878] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0879] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0880] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0881] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0882] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0883] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0884] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0885] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0886] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0887] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0888] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0889] The present invention relates to a system for displaying real-time scoring results during sports viewing. The system receives a video feed, analyzes it, generates technical scores, composition scores, and deductions, and visually displays the results. Specific embodiments of the system are described below.
[0890] System Overview
[0891] server
[0892] First, the server receives a video feed from a camera or a video file. The video feed contains footage of the athletes' performances and is processed frame by frame. The server initializes the video capture device and continuously receives the video data.
[0893] The server then extracts each frame from the received video feed and passes it to an AI evaluation module, which analyzes each frame and generates a technical score, composition score, and deductions for mistakes. This analysis result is obtained for each frame, allowing for real-time evaluation.
[0894] User terminal
[0895] The user's device receives the analysis results sent from the server and displays them as an overlay on the video feed frame. The numerical values for technical scores, composition scores, and deductions are displayed as text in the appropriate position within the frame. This allows the audience to visually check the evaluation score for each performance in real time, and intuitively understand the scoring criteria.
[0896] Program processing
[0897] server
[0898] 1. The server continues to receive the video feed from the video capture device.
[0899] 2. The server extracts each frame from the received video feed.
[0900] 3. The server passes each extracted frame to the AI evaluation module.
[0901] 4. The AI evaluation module analyzes each frame and generates an evaluation result including technical points, composition points, and deductions.
[0902] User terminal
[0903] 1. The user device receives the analysis results sent from the server.
[0904] 2. The user device overlays the analysis results on frames of the video feed.
[0905] 3. The user's device will display the numerical values for technical score, composition score, and deduction points for each frame in text at the appropriate position.
[0906] Specific examples
[0907] For example, in a figure skating competition, the server receives a frame from the video feed in which an athlete performs a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points to this jump, and a technical score of 7.8 points to the subsequent spin. The user device receives these evaluation results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can see this in real time and understand the performance evaluation on the spot. This improves the transparency and enjoyment of watching sports.
[0908] The processing flow will be explained below.
[0909] Step 1:
[0910] The server initializes the video capture device to receive the video feed from the camera or video file. A video capture object is created using cv2.VideoCapture(0), which starts capturing real-time video data.
[0911] Step 2:
[0912] The server reads the video data frame by frame from the received video feed, specifically by using the read method of the video capture object to extract the frame-by-frame images.
[0913] Step 3:
[0914] The server passes the acquired frames to the AI evaluation module, which receives the images frame by frame as preprocessing and extracts features related to the player's movements from the input.
[0915] Step 4:
[0916] The server receives the analysis results from the AI evaluation module, which calculates the technical score, composition score, and deduction score for each frame and returns the evaluation results including this data.
[0917] Step 5:
[0918] The server formats the evaluation results and converts them into a data format for sending to the user's device. Specifically, the evaluation results are converted into a flexible data format such as JSON.
[0919] Step 6:
[0920] The user terminal receives the analysis result data sent from the server, analyzes this data, and extracts the necessary information.
[0921] Step 7:
[0922] The user device overlays the analysis results on the video feed frame, using the cv2.putText method to display numerical information such as technical scores, composition scores, and deductions in the appropriate positions on the frame.
[0923] Step 8:
[0924] The user device displays the frame containing the overlay information on the screen in real time using the cv2.imshow method, allowing the audience to instantly see the evaluation results on the screen.
[0925] Specific examples
[0926] In a figure skating competition, the server receives the frame at the moment a triple axel jump is performed. The AI evaluation module analyzes this frame and assigns a technical score of 9.5 points for the jump. It also analyzes the frame for the subsequent spin and assigns a technical score of 7.8 points. These evaluation results are sent from the server to the user's device, which receives them. The user's device displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can view this in real time and instantly understand the evaluation of each performance. This process improves the transparency of scoring criteria and deepens understanding of sports spectators.
[0927] Example 1
[0928] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0929] Conventional sports viewing systems make it difficult for spectators to visually understand the technical and compositional scores of athletes in real time. Furthermore, the lack of transparency and speed in the evaluation process reduces the enjoyment of watching the game.
[0930] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0931] In this invention, the server includes means for receiving a video feed from a video capture device, means for extracting frames from the video feed and analyzing the video data to generate evaluation results regarding technical scores, composition scores, and deductions, means for overlaying the evaluation results on the frames in real time, and means for visually displaying the evaluation results in text format, thereby enabling spectators to visually check the evaluation scores in real time and enjoy fast and transparent evaluation.
[0932] "Video capture device" means equipment or software for receiving a video feed, including cameras and video streaming devices.
[0933] A "video feed" is a continuously distributed stream of moving image data, including footage of sporting matches and events.
[0934] "Frames" are the individual still images that make up a video feed and are played back in succession to form a moving image.
[0935] "Motion image data" is continuous image data consisting of multiple frames, and is used to analyze motion and movement.
[0936] "Technical score" is a technical evaluation score of an athlete's performance, based on the degree of perfection and difficulty of a particular movement.
[0937] "Composition score" is an evaluation score based on the composition and direction of the athlete's entire performance, taking into account creativity and expressiveness.
[0938] "Deductions" are points deducted from an athlete's technical and compositional scores for mistakes or incorrect movements during their performance.
[0939] The "evaluation result" is an analysis result including technical points, composition points, and deductions, and is generated for each frame.
[0940] "Overlay display" refers to the process of displaying the evaluation results overlaid on frames of the video feed, allowing the evaluation to be visually confirmed in real time.
[0941] "Text format" is a format in which information is displayed as text, and presents numerical values, evaluation points, etc. in a visually easy-to-understand manner.
[0942] A "deep learning evaluation module" is a module that uses artificial intelligence technology to analyze video data and evaluate technical and compositional aspects, and includes neural networks.
[0943] The present invention relates to a system for displaying real-time scoring results during sports viewing. The system receives video feeds from a video capture device, analyzes the video feeds, evaluates technical scores, composition scores, and deductions, and visually displays the evaluation results. Specific embodiments of the system are described below.
[0944] server
[0945] First, the server initializes a video capture device (e.g., HD camera, IP camera) as a video acquisition device and receives the video feed. The video feed contains footage of the athletes' performances, which is processed frame by frame. The server extracts the frames using an image processing library such as OpenCV.
[0946] Next, the server inputs each extracted frame into a deep learning evaluation module (e.g., a model using TensorFlow or PyTorch). This deep learning evaluation module analyzes each frame and evaluates technical points, composition points, and deductions. Analysis results are generated for each frame, allowing for real-time evaluation.
[0947] Specifically, the server receives a frame of a figure skating competition in which a skater performs a triple axel jump. This frame is preprocessed and sent to the evaluation module for technical evaluation. The evaluation module assigns a technical score of 9.5 points to the jump and 7.8 points to the subsequent spin.
[0948] User terminal
[0949] The user device receives the analysis results sent from the server. This data is sent and received using a real-time communication protocol such as WebSocket. Based on the received analysis results, the user device displays the evaluation results on the video feed frames. Specifically, numerical information such as technical scores, composition scores, and deductions is overlaid on the video frames using the HTML5 Canvas element and OpenGL.
[0950] Furthermore, the user's device will display the evaluation score in text format at the appropriate position for each frame. For example, text such as "Triple Axel: 9.5 points" or "Spin: 7.8 points" will be displayed in the bottom left or bottom right. This allows the audience to visually check the evaluation score in real time.
[0951] This system will improve the transparency of sports viewing and increase the enjoyment of watching, and will also enable spectators to intuitively understand the evaluation of performances in real time.
[0952] Prompt Sentence Examples
[0953] Below are some example prompts to input to the generative AI model:
[0954] "Describe an AI system that analyzes a video feed of figure skating performances and displays the technical and compositional scores for each performance in real time."
[0955] "Please tell us the specific technical details of a sports viewing system that displays the types of moves performed by athletes and their evaluation scores in real time."
[0956] The above is a specific embodiment for carrying out the present invention.
[0957] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0958] Step 1:
[0959] The server receives a video feed from a video capture device. The video capture device can be an HD camera or an IP camera. The input is the video feed from the video capture device, and the output is video image data divided into frames. This video feed is continuously captured and stored in the server's memory.
[0960] Step 2:
[0961] The server extracts individual frames from the video feed. It uses an image processing library such as OpenCV to acquire each frame. The input is the video feed received in step 1, and the output is image data for each frame. Specifically, it uses the OpenCV read function to extract the frames.
[0962] Step 3:
[0963] The server passes each extracted frame to a deep learning evaluation module. Using a deep learning framework such as TensorFlow or PyTorch, the module extracts and evaluates the frame's features. The input is the frame image extracted in step 2, and the output is the evaluation results, including technical scores, composition scores, and deductions. A convolutional neural network (CNN) is used for this evaluation to analyze the movement.
[0964] Step 4:
[0965] The deep learning evaluation module analyzes each frame and generates evaluation results in terms of technical points, composition points, and deductions. Specifically, it uses CNN to identify actions within the frame and calculate a score for each action. The input here is the frame image passed in step 3, and the output is the evaluation result for each frame.
[0966] Step 5:
[0967] The server sends the generated evaluation results to the user's device in real time using WebSocket. The input is the evaluation results generated in step 4, and the output is the analysis data sent to the user's device.
[0968] Step 6:
[0969] The user terminal receives the analysis results sent from the server. The input here is the evaluation results sent in step 5, and the output is the evaluation data stored in the memory on the user terminal.
[0970] Step 7:
[0971] The user device overlays the received analysis results on the video feed frame. Using the HTML5 Canvas element or OpenGL, the numerical information of the evaluation results is displayed overlaid on the video frame. The input is the evaluation results received in step 6, and the output is the overlaid video frame. Specific operations include displaying text in the appropriate position.
[0972] Step 8:
[0973] The user device displays the technical score, composition score, and deduction scores in text format at the appropriate position on the video frame. The input is the frame overlaid in step 7, and the output is real-time video that allows the audience to visually check the evaluation scores. For example, text information such as "Triple Axel: 9.5 points" or "Spin: 7.8 points" is displayed.
[0974] This series of processes allows spectators to visually check the evaluation scores for athletes' performances in real time, improving the transparency and enjoyment of watching sports.
[0975] (Application example 1)
[0976] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0977] With conventional sports viewing systems, it was difficult for spectators to check the evaluation of performances and plays in real time. In particular, there were limited means for them to visually understand the evaluation of technical points, composition points, and deductions. This reduced the transparency and enjoyment of the viewing experience, and prevented spectators from fully experiencing the appeal of sports. Another issue was the lack of systems that took into account use on mobile devices such as smartphones.
[0978] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0979] In this invention, the server includes means for receiving a video feed, means for analyzing moving image data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for overlaying and displaying the evaluation results on the frames, means for receiving a video feed from a live streaming server, means for analyzing the video frames with an AI evaluation module and generating evaluation results, and means for overlaying and displaying the analysis results on the video feed in real time, thereby enabling spectators to visually understand the evaluation results of sports performances and plays being live-streamed in real time using their smartphones.
[0980] A "video feed" is a continuous stream of image data transmitted from a camera or other video capture device.
[0981] A "frame" is an individual still image in a video feed that is displayed in succession to form a moving image.
[0982] "Motion image data" refers to video data consisting of a series of frames obtained from a video feed.
[0983] "Technical points" are the points awarded to athletes in sports competitions based on their technical performance and skills.
[0984] "Composition points" are points given in sports competitions to evaluate the overall composition and artistry of an athlete's performance.
[0985] "Deductions" are points that are deducted in sports competitions when a player commits a mistake or acts in violation of the rules.
[0986] "AI Evaluation Module" means a software module that uses artificial intelligence technology to analyze video frames and generate technical scores, composition scores, and deduction scores.
[0987] A "live streaming server" is a server for delivering video feeds in real time over the Internet.
[0988] "Overlay display" is a technique for displaying additional information superimposed on a video frame.
[0989] "Real-time" means processed and displayed immediately, without delay or lag.
[0990] A "smartphone" is a portable information terminal that has mobile phone functions and can be used by installing additional applications.
[0991] "Analysis results" are the evaluation data of technical points, composition points, and deduction points generated by the AI evaluation module.
[0992] "Visual presentation" means displaying information in a way that is easy for the audience to see, and is a technique that allows information to be intuitively understood.
[0993] This invention relates to a system for displaying evaluation results in real time while watching sports. The system can receive a video feed, analyze it, generate evaluations of technical points, composition points, and deductions, and display the results visually.
[0994] System Overview
[0995] server
[0996] The server receives the video feed from the live streaming server and extracts video data from the video feed frame by frame. It initializes the video capture device and continuously receives video data. The server then extracts each frame from the received video feed and passes it to the AI evaluation module. The AI evaluation module analyzes each frame and generates technical scores, composition scores, and deductions for mistakes. The analysis results are obtained for each frame, allowing for real-time evaluation.
[0997] User terminal
[0998] The user's device receives the analysis results sent from the server and displays them as an overlay on the video feed frame. The numerical values for technical points, composition points, and deductions are displayed as text in the appropriate position within the frame. This allows spectators to visually check the evaluation scores for each performance and play in real time, and to intuitively understand the evaluation criteria.
[0999] Program processing
[1000] The implementation of this system requires several key software and hardware components: OpenCV is used to sample the video feed and extract frames, TensorFlow is used to generate scores for the AI evaluation module, and the phone's native UI framework (e.g., UIKit for iOS, Jetpack Compose for Android) is used to render the overlay display.
[1001] Specific hardware and software
[1002] Smartphone: iOS or Android device
[1003] Live Streaming Server: Streams video feeds using ffmpeg
[1004] AI evaluation module: Machine learning libraries such as TensorFlow
[1005] Frame analysis software: OpenCV
[1006] UI Framework: UIKit or Jetpack Compose
[1007] Specific examples
[1008] For example, in a figure skating competition, the server receives a frame from the video feed of the moment an athlete performs a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points for this jump and a technical score of 7.8 points for the subsequent spin. The user's device receives these evaluation results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." Spectators can see this in real time and understand the performance evaluation on the spot. This improves the transparency and enjoyment of watching sports.
[1009] Prompt Sentence Examples
[1010] Create an application that displays real-time scoring results while watching sports. The app receives a video feed from a live streaming server and analyzes each frame with a TensorFlow model to generate and display technical scores, composition scores, and deductions. The software used is ffmpeg, OpenCV, and TensorFlow.
[1011] By implementing this invention, spectators can use their smartphones to visually understand the evaluation results of live-streamed sports performances and plays in real time.
[1012] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1013] Step 1:
[1014] The server receives a video feed from the live streaming server. It initializes a video capture device and continuously receives video data. The input is the video feed from the live streaming server, and the output is a stream of video frames. Specifically, the server initializes a video capture device and continuously processes the received video stream.
[1015] Step 2:
[1016] The server extracts each frame from the received video feed. It uses OpenCV to extract the frames and convert them into individual still image data. The input is the video feed, and the output is still image data for each frame. Specifically, it uses OpenCV functions to extract still images for each frame from the video stream.
[1017] Step 3:
[1018] The server passes each extracted frame to an AI evaluation module, which uses TensorFlow to analyze the frames and generate evaluations for technical, compositional, and deduction points. The input is still image data for each frame, and the output is the evaluation results for technical, compositional, and deduction points. Specifically, a pre-trained model is used to analyze the data for each frame and calculate the points.
[1019] Step 4:
[1020] The server transmits the evaluation results generated by the AI evaluation module to the user terminal. The evaluation results are transmitted in real time. The input is the evaluation result for each frame, and the output is the evaluation data transmitted over the network. Specifically, the server converts the evaluation results into an appropriate format and transmits them over the network.
[1021] Step 5:
[1022] The user device overlays the analysis results received from the server on the video feed frame. The input is the evaluation results sent from the server, and the output is the overlaid video frame. Specifically, the device's native UI framework is used to display the evaluation results in text on the frame in real time.
[1023] Step 6:
[1024] The user checks the evaluation results in real time using a smartphone. The input is the video feed displayed on the user's device, and the output is the user's visual understanding of the information. Specifically, the user checks the performances and plays in the video stream along with the evaluation scores overlaid in real time.
[1025] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1026] The present invention relates to a system that displays real-time scoring results during sports viewing and customizes the display content by recognizing the user's emotions. The system receives a video feed, analyzes it, generates technical scores, composition scores, and deductions, and visually displays the results. It also incorporates an emotion engine that recognizes the user's emotions. Specific embodiments of the system are described below.
[1027] System Overview
[1028] server
[1029] First, the server receives a video feed from a camera or a video file. The video feed contains footage of the athletes' performances and is processed frame by frame. The server initializes the video capture device and continuously receives the video data.
[1030] The server then extracts each frame from the received video feed and passes it to an AI evaluation module, which analyzes each frame and generates a technical score, composition score, and deductions for mistakes. This analysis result is obtained for each frame, allowing for real-time evaluation.
[1031] Furthermore, the server is equipped with an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's facial expressions and voice to identify emotions such as enjoyment, surprise, or dissatisfaction. Based on the recognized emotion, the server customizes the display of the evaluation results. For example, if the user shows interest, it can provide additional detailed information.
[1032] User terminal
[1033] The user's device receives the analysis results sent from the server as well as information from the emotion engine. Based on this, an overlay display is performed on the video feed frame. As text, numerical values for technical points, composition points, and deductions are displayed in appropriate positions within the frame. Furthermore, a UI (user interface) can be constructed according to the user's emotions, and the displayed content can be dynamically changed.
[1034] Program processing
[1035] server
[1036] 1. The server continues to receive the video feed from the video capture device.
[1037] 2. The server extracts each frame from the received video feed.
[1038] 3. The server passes each extracted frame to the AI evaluation module.
[1039] 4. The AI evaluation module analyzes each frame and generates an evaluation result including technical points, composition points, and deductions.
[1040] 5. The server uses an emotion engine to recognize emotions from the user's face and voice.
[1041] 6. Customize the display of rating results based on the emotions recognized by the emotion engine.
[1042] User terminal
[1043] 1. The user device receives the analysis results and emotion engine information sent from the server.
[1044] 2. The user device overlays the analysis results and emotion data onto frames of the video feed.
[1045] 3. The user's device will display the numerical values for technical score, composition score, and deduction points for each frame in text at the appropriate position.
[1046] 4. Dynamically change the UI based on the user's emotions, for example highlighting important information if the user is surprised.
[1047] Specific examples
[1048] For example, in a figure skating competition, the server receives a frame from the video feed of a skater performing a triple axel jump. The AI evaluation module assigns a technical score of 9.5 points for the jump and a technical score of 7.8 points for the subsequent spin. These evaluation results are then sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice and recognizes their surprise. The user's device receives the results and displays a text overlay on the frame, reading "Triple axel: 9.5 points" and "Spin: 7.8 points." To further emphasize their surprise, the user's device also displays a comment such as "Great jump!" Spectators can view this in real time and receive interactive feedback based on the evaluation of each performance and their own emotions. This process increases the transparency of scoring criteria and further enhances the enjoyment of watching sports.
[1049] The processing flow will be explained below.
[1050] Step 1:
[1051] The server initializes a video capture device to receive the video feed from a camera or a video file. Specifically, it uses cv2.VideoCapture(0) to get real-time video data from the device.
[1052] Step 2:
[1053] The server extracts individual frames from the received video feed, using the read method of the video capture object to retrieve the image data for each frame.
[1054] Step 3:
[1055] The server passes each extracted frame to an AI evaluation module, which uses a pre-trained model to analyze the content of each frame and calculate technical and compositional scores, as well as deductions.
[1056] Step 4:
[1057] The server receives the analysis results from the AI evaluation module and converts them into a data format (e.g., JSON format), which is then prepared for transmission to the user's device.
[1058] Step 5:
[1059] The server uses an emotion engine to analyze the user's facial expressions and voice in real time to recognize their emotions. For example, it analyzes camera footage and microphone input to determine whether the user is enjoying, surprised, or dissatisfied.
[1060] Step 6:
[1061] The server customizes the displayed evaluation results based on the user's emotions. For example, if the user is surprised, it highlights specific technical or structural features.
[1062] Step 7:
[1063] The server then sends the customized evaluation results to the user's device, using a network to transfer data in real time.
[1064] Step 8:
[1065] The user device receives the analysis result data and emotion data sent from the server, and prepares an overlay display on the frame of the video feed based on this data.
[1066] Step 9:
[1067] The user's device overlays the analysis results on the frame. Specifically, the cv2.putText method is used to display the numerical information on technical points, composition points, and deductions in text format in the appropriate location.
[1068] Step 10:
[1069] The user device dynamically changes the UI based on the user's emotions. For example, if it detects that the user is surprised, it will display a comment such as "Great jump!"
[1070] Step 11:
[1071] The user device displays frames containing overlay information on the screen in real time, allowing the audience to see the evaluation results and interactive feedback in real time.
[1072] Specific examples
[1073] In a figure skating competition, the server receives a frame from the video feed of the moment an athlete performs a triple axel jump. The AI evaluation module analyzes this frame and assigns a technical score of 9.5 points for the jump. It also analyzes the frame for the spin and assigns a technical score of 7.8 points. These evaluation results are sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice to recognize surprise. The user's device receives this and overlays text such as "Triple axel: 9.5 points" and "Spin: 7.8 points" on the frame. It also displays a comment such as "Great jump!" to emphasize surprise. This allows spectators to receive interactive feedback in real time based on the evaluation of each performance and their own emotions, improving the transparency and enjoyment of sports viewing.
[1074] Example 2
[1075] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1076] Conventional sports viewing systems have the problem that it is difficult to visually display the performance evaluation of athletes in real time, and they do not provide interactive feedback according to the user's emotions. As a result, spectators cannot immediately grasp the detailed evaluation results during the performance, which limits the viewing experience.
[1077] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1078] In this invention, the server includes means for receiving a video feed, means for analyzing video data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for recognizing a user's emotions and customizing the display content, and means for overlaying the evaluation results on the frames, thereby making it possible to accurately evaluate a player's performance in real time, display the evaluation results, and provide interactive feedback according to the user's emotions.
[1079] "Video Feed" refers to a continuous stream of real-time or recorded video data.
[1080] "Motion image data" refers to digital data that contains image information that changes over time, and is typically made up of a collection of individual still images called frames.
[1081] "Technical score" is a numerical value that indicates the evaluation of the technical elements of an athlete's performance, including the accuracy and difficulty of technical movements such as jumps and spins.
[1082] "Composition score" is a numerical value that indicates an evaluation of the composition and artistic quality of a performance, and includes, for example, expressiveness and the flow of the performance.
[1083] "Deductions" are points that are subtracted from the technical or composition score due to a mistake or violation of rules.
[1084] "User emotion" refers to the subjective emotional state, such as enjoyment, surprise, or dissatisfaction, that a spectator or user feels while watching a sporting event.
[1085] "Customizing the display content" refers to dynamically changing the information and interface content displayed in response to the user's emotions.
[1086] "Overlay display" refers to a method of displaying evaluation results or additional information overlaid on the video feed.
[1087] This invention relates to a system that displays the scores of sports games in real time and customizes the display content by recognizing the emotions of users. The system is composed of a server and a user terminal.
[1088] First, the server initializes the video capture device and receives the athletes' performances in real time. Specifically, it uses OpenCV to acquire video data from IP cameras and USB cameras. The received video feed is extracted frame by frame and passed to the AI evaluation module. The AI evaluation module uses TensorFlow and PyTorch to analyze each frame and generate technical scores, composition scores, and deductions. This makes it possible to evaluate the athletes' performances in real time.
[1089] The server is also equipped with an emotion engine that recognizes the user's emotions. Using services such as Google Cloud Natural Language API and IBM Watson, it analyzes the user's facial expressions and voice to identify emotions such as whether the user is enjoying, surprised, or dissatisfied. After identifying the emotion, the server customizes the display of the evaluation results. For example, if the user is surprised, it can add a comment such as "Great jump!"
[1090] Next, the user device receives the analysis results and emotion engine information sent from the server. <canvas>Using JavaScript and JavaScript, the analysis results are overlaid on frames of the video feed. By visually displaying numerical values for technical, compositional, and deduction scores, the audience can see a detailed evaluation of each performance in real time. The UI also dynamically changes based on the user's emotions, highlighting important information as needed.
[1091] Take a figure skating competition as an example. The server receives the moment of a triple axel jump from a video feed, and the AI evaluation module assigns a technical score of 9.5 points for this jump and 7.8 points for the following spin. These evaluation results are sent to the user's device in real time, and the emotion engine recognizes that the user is surprised. The user's device receives this and displays a text overlay saying "Triple axel: 9.5 points" and "Spin: 7.8 points," along with the comment "Great jump!"
[1092] Example prompt sentence:
[1093] "Analyze a figure skater's triple axel jump in real time and rate it for technical and compositional points. Highlight results that amaze the user."
[1094] As described above, this invention allows users to check player evaluations in real time and receive interactive feedback based on their own emotions. This system will further enhance the enjoyment of watching sports.
[1095] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1096] Step 1:
[1097] The server initializes the video capture device and receives the athletes' performances in real time. Specifically, the server acquires video data using OpenCV's cv2.VideoCapture. The input is video data from the video capture device, and the output is video data divided into individual frames.
[1098] Step 2:
[1099] The server extracts frames one by one from the received video feed. Specifically, the server reads frames continuously using the OpenCV read() method and adds them to a list. The input is continuous video data from the video capture device, and the output is video data frame by frame.
[1100] Step 3:
[1101] The server passes each extracted frame to the AI evaluation module. Specifically, the server inputs the frame data into the AI evaluation module using TensorFlow or PyTorch. The input is video data for each frame, and the output is analyzed evaluation data (technical score, composition score, and deductions).
[1102] Step 4:
[1103] The AI evaluation module analyzes each frame and generates technical scores, composition scores, and deductions. Specifically, the AI model analyzes specific movements within a frame and outputs evaluation data such as "jump: 9.5 points" or "spin: 7.8 points." The input is the video data for each frame, and the output is evaluation data for each performance.
[1104] Step 5:
[1105] The server uses an emotion engine to recognize emotions from the user's face and voice. Specifically, the server analyzes the user's facial expression and voice data using Google Cloud Natural Language API and IBM Watson. The input is the user's facial expression and voice data, and the output is the user's emotional state (enjoyed, surprised, dissatisfied, etc.).
[1106] Step 6:
[1107] The display content is customized based on the emotions recognized by the emotion engine. Specifically, the server generates additional information or comments to attract the user's attention based on the analysis results. For example, if the user is surprised, it adds a comment such as "Great jump!" The input is the user's emotional state and evaluation data for each frame, and the output is customized display data.
[1108] Step 7:
[1109] The user device receives the analysis results and emotion engine information sent from the server. Specifically, the user device communicates with the server using WebSocket or RESTful API to receive data. The input is the analysis results and emotion information sent from the server, and the output is data stored on the user device.
[1110] Step 8:
[1111] The user device overlays the analysis results and emotion data on the video feed frame. <canvas>It uses elements and JavaScript to overlay analytics data on a video feed. The input is analytics results and emotion data, and the output is a visual display on the user's device.
[1112] Step 9:
[1113] The user's device displays the technical score, composition score, and deduction score values for each frame in the appropriate position. Specifically, it performs coordinate calculations to display the analysis data as text in the appropriate position within the frame. The input is the evaluation data for each frame, and the output is a text overlay display for each frame.
[1114] Step 10:
[1115] The user device dynamically changes the UI according to the user's emotions. Specifically, it uses CSS and JavaScript to change UI elements and highlight important information. For example, if the user is surprised, it changes the text color or highlights a specific comment. The input is the user's emotional data and analysis results, and the output is a dynamically changed UI display.
[1116] (Application example 2)
[1117] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1118] In conventional sports viewing systems, it was difficult for spectators to know the evaluation results of athletes' technical and compositional scores in real time, and the evaluation criteria lacked transparency. Furthermore, there was no way to provide a more interactive and engaging viewing experience by incorporating spectators' emotions and excitement. As a result, it was difficult to increase spectator satisfaction.
[1119] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1120] In this invention, the server includes means for receiving a video feed, means for analyzing video data for each frame from the video feed and generating evaluation results regarding technical scores, composition scores, and deductions, means for overlaying the evaluation results on the frames, and means for analyzing a user's emotions and customizing the display content of the evaluation results. This allows viewers to know the evaluation results, such as technical scores and composition scores, in real time when watching a sports game, and further provides an interactive viewing experience that responds to the viewers' emotions.
[1121] A "video feed" is a continuous video signal coming from a camera or other video capture device.
[1122] A "frame" refers to an individual still image within a video feed, with a series of frames forming a moving image.
[1123] "Video data" refers to the digital video data that comprises each frame in a video feed.
[1124] "Technical score" is a numerical evaluation of an athlete's technical performance in a sporting event.
[1125] "Composition points" are a numerical evaluation of the composition and expressiveness of a performance in a sporting event.
[1126] "Deduction points" refers to the evaluation of a sporting event in which negative points are given for technical errors or violations of rules.
[1127] "Evaluation results" refers to the overall score including technical points, composition points, and deductions.
[1128] "Overlay display" refers to a technique for displaying additional information overlaid on top of a video frame.
[1129] "User's emotion" refers to the subjective emotional state that the user feels while watching a sporting event, and is analyzed from facial expressions and voice.
[1130] "Emotion engine" refers to a software engine that recognizes and analyzes emotions from the user's face and voice.
[1131] "UI" stands for user interface and refers to the screen layout and design that allows users to interact with systems and applications.
[1132] "Real-time" refers to information and data processing occurring in real time, with results reflected almost immediately.
[1133] A system for implementing this invention consists of a server and a user terminal. The server receives a video feed and analyzes the video data frame by frame to generate evaluation results in terms of technical points, composition points, and deductions. It also analyzes the user's emotions and customizes the display of the evaluation results. The user terminal displays the evaluation results as an overlay, visually presenting them to the user.
[1134] Specifically, the server receives live feeds from cameras and video capture devices. This live feed contains the athletes' performances and analyzes them frame by frame. An AI evaluation module using TensorFlow is used for the analysis, and evaluation results are generated in real time, including technical scores, compositional scores, and deductions. These evaluation results are then sent to the user's device.
[1135] The server is equipped with an emotion engine that recognizes emotions from viewers' faces and voices. This emotion engine uses VaderSentiment and TensorFlow emotion recognition models. Viewers' facial expressions and voice data are analyzed using the dlib library to grasp the user's emotional state in real time. As a result, the display of the evaluation results is customized, such as highlighting important information if the user is surprised.
[1136] The user's device receives the analysis results and emotion data sent from the server. A Python-based GUI application is installed on the user's device, which overlays the live feed frames with information on technical and compositional points, as well as deductions. Furthermore, the application dynamically displays comments and supplementary information based on the user's emotions.
[1137] As a concrete example, in a figure skating competition, the server analyzes the frame at the moment when the skater performs a triple axel jump. The AI evaluation module generates a technical score of 9.5 points for this jump, and a technical score of 7.8 points for the subsequent spin. These evaluation results are sent from the server to the user's device. At the same time, the emotion engine analyzes the user's facial expressions and voice and recognizes that the user is surprised. The user's device receives this and overlays the frame with text such as "Triple axel: 9.5 points" and "Spin: 7.8 points," along with additional comments such as "Great jump!"
[1138] An example prompt is:
[1139] "Design a system that displays technical and compositional scores in real time based on viewer emotions during live-streamed sporting events. Utilize an AI evaluation module and emotion recognition engine to dynamically display cheering comments, such as 'Great jump!', when a viewer expresses surprise."
[1140] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1141] Step 1:
[1142] The server receives the video feed from the camera or video capture device. Here, the input is a live video feed, and each frame of the video feed is continuously sent to the server. This step captures the athlete's performance in real time. The output is a continuous video signal.
[1143] Step 2:
[1144] The server extracts each frame from the received video feed. The input is the video feed, and data is extracted frame by frame. Here, the actual processing is performed to extract the video data obtained from the video capture device one frame at a time. The output is frame data, which are individual still images.
[1145] Step 3:
[1146] The server passes the extracted frame data to the AI evaluation module. The input is the frame data, which is analyzed frame by frame using TensorFlow. The module calculates the technical score, composition score, and deduction points, and outputs the evaluation results as numerical data.
[1147] Step 4:
[1148] The server is equipped with an emotion engine that analyzes emotions from the user's facial expressions and voice. The input is the user's facial expression data and voice data, which are analyzed using the dlib library and VaderSentiment. In this step, specific processing is performed to identify the user's emotional state, such as whether they are surprised, amused, or frustrated. The output is the user's emotional state data.
[1149] Step 5:
[1150] The server combines the evaluation results and the user's emotional state to customize the content of the evaluation results. The input is the numerical evaluation results and emotional state data, and the display content is dynamically changed based on these. In this step, specific operations are performed, such as adding a comment such as "Great jump!" if the user is surprised. The output is customized evaluation result data.
[1151] Step 6:
[1152] The server sends the customized evaluation result data to the user terminal. The input is the customized evaluation result data, which is sent to the user terminal via the network. In this step, the specific operation of data transfer is performed. The output is the evaluation result data sent to the user terminal.
[1153] Step 7:
[1154] The user device uses the received evaluation result data and emotion data to overlay a frame of the video feed. The input is the customized evaluation result data and emotion data, and specific processing is performed to display text and comments based on these. The output is a visual display result.
[1155] Step 8:
[1156] The user device dynamically changes the UI according to the user's emotional response. The input is real-time emotional data, and specific actions are taken to change the interface and presented information based on this. The output is an updated UI.
[1157] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1158] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1159] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1160] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1161] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1162] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1163] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1164] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1165] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1166] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1167] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1168] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1169] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1170] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1171] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1172] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1173] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1174] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1175] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1176] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1177] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1178] The following is further disclosed regarding the above embodiment.
[1179] (Claim 1)
[1180] means for receiving a video feed;
[1181] means for analyzing the video data frame by frame from the video feed and generating evaluation results in terms of technical points, composition points, and deduction points;
[1182] means for overlaying and displaying the evaluation results on the frame;
[1183] A system including:
[1184] (Claim 2)
[1185] 10. The system of claim 1, configured to analyze and display the evaluation results in real time.
[1186] (Claim 3)
[1187] 2. The system according to claim 1, further comprising means for displaying the numerical values of the evaluation results in text on the frame in order to visually present the details of the analysis results to the audience.
[1188] "Example 1"
[1189] (Claim 1)
[1190] means for receiving a video feed from a video capture device;
[1191] means for extracting frames from the video feed and analyzing the video data to generate evaluation results in terms of technical points, composition points, and deduction points;
[1192] means for overlaying and displaying the evaluation results on the frame in real time;
[1193] means for visually displaying the evaluation results in text format;
[1194] A system including:
[1195] (Claim 2)
[1196] 10. The system of claim 1, wherein the system is configured to use a deep learning evaluation module to analyze the characteristics of each frame and perform a technical evaluation of the motion when the evaluation result is generated.
[1197] (Claim 3)
[1198] 2. The system according to claim 1, further comprising means for displaying the numerical values of the evaluation results in text format at an appropriate position on the frame in order to visually present the details of the analysis results to the audience.
[1199] "Application Example 1"
[1200] (Claim 1)
[1201] means for receiving a video feed;
[1202] means for analyzing the video data frame by frame from the video feed and generating evaluation results in terms of technical points, composition points, and deduction points;
[1203] means for overlaying and displaying the evaluation results on the frame;
[1204] means for receiving a video feed from a live streaming server;
[1205] means for analyzing the video frames by an AI evaluation module and generating an evaluation result;
[1206] A means to overlay analysis results on the video feed in real time;
[1207] A system including:
[1208] (Claim 2)
[1209] 10. The system of claim 1, configured to analyze and display evaluation results in real time and operating on a smartphone.
[1210] (Claim 3)
[1211] 2. The system according to claim 1, further comprising means for displaying the numerical values of the evaluation results in text on the frame in order to visually present the details of the analysis results to the audience.
[1212] "Example 2: Combining Emotion Engines"
[1213] (Claim 1)
[1214] means for receiving a video feed;
[1215] means for analyzing the video data frame by frame from the video feed and generating evaluation results in terms of technical points, composition points, and deduction points;
[1216] means for recognizing a user's emotions and customizing the displayed content;
[1217] means for overlaying and displaying the evaluation results on the frame;
[1218] A system including:
[1219] (Claim 2)
[1220] 10. The system of claim 1, configured to analyze and display the evaluation results in real time.
[1221] (Claim 3)
[1222] 2. The system according to claim 1, further comprising means for displaying the numerical values of the evaluation results in text on the frame in order to visually present the details of the analysis results to the audience.
[1223] "Application example 2 when combining emotion engines"
[1224] (Claim 1)
[1225] means for receiving a video feed;
[1226] means for analyzing the video data frame by frame from the video feed and generating evaluation results in terms of technical points, composition points, and deduction points;
[1227] means for overlaying and displaying the evaluation results on the frame;
[1228] means for analyzing the user's emotions and customizing the display content of the evaluation results;
[1229] A system including:
[1230] (Claim 2)
[1231] 10. The system of claim 1, configured to analyze and display the evaluation results in real time.
[1232] (Claim 3)
[1233] 2. The system according to claim 1, further comprising means for displaying the numerical values of the evaluation results in text on the frame in order to visually present the details of the analysis results to the audience.
[1234] (Claim 4)
[1235] 2. The system according to claim 1, wherein the system uses a user emotion analysis engine to recognize the user's emotions from facial expressions and voice, and dynamically changes the UI according to the displayed content of the evaluation results.
[1236] (Claim 5)
[1237] 5. The system of claim 4, further comprising means for highlighting important information when the user is surprised. [Explanation of symbols]
[1238] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / canvas> < / canvas> < / url:> < / canvas> < / canvas> < / url:> < / canvas> < / canvas> < / url:> < / canvas> < / canvas>
Claims
1. means for receiving a video feed; means for analyzing the video data frame by frame from the video feed and generating evaluation results in terms of technical points, composition points, and deduction points; means for overlaying and displaying the evaluation results on the frame; A system including:
2. 10. The system of claim 1, configured to analyze and display the evaluation results in real time.
3. 2. The system according to claim 1, further comprising means for displaying the numerical values of the evaluation results in text on said frame in order to visually present the details of the analysis results to an audience.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A