System
The system automates amateur sports game footage analysis by extracting features and updating a learning model, reducing user effort and improving accuracy over time.
Patent Information
- Application Number
- JP2024137099
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Current analysis methods for amateur sports game footage require significant time and effort, struggle with automatic extraction of specific scenes, and are inaccurate due to variations in shooting distance and angle.
A system that automates the analysis of amateur sports game footage by extracting features such as player movements and ball trajectory, allowing users to verify and update a learning model for improved accuracy.
The system significantly reduces user workload and continuously improves analysis accuracy through incremental learning, enabling efficient and accurate extraction of intended scenes.
Smart Images

Figure 2026033978000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] While analysis of game footage has become commonplace in amateur sports, current analysis methods require significant time and effort, making them particularly challenging for sports with mixed offensive and defensive play. Furthermore, automatic extraction of specific scenes is difficult, and extracting scenes consistent with the user's intent requires significant effort. Furthermore, analysis of footage shot at different distances and angles is difficult, negatively impacting the accuracy and efficiency of analysis. This invention aims to solve these problems by automating and streamlining the analysis of amateur sports game footage. [Means for solving the problem]
[0005] The present invention provides a means for receiving video of a scene to be learned and extracting features from the video. It also provides a means for receiving game video and extracting scenes from the game video that match the features. It also provides a means for presenting the extracted scenes to a user, allowing the user to determine whether the scene is a predetermined scene, recording the user's determination, and updating the learning model to improve extraction accuracy in future. The system may further include a means for extracting player movements, positioning, and ball trajectory as features of the scene to be learned. It may also include a means for extracting features from video shot at different distances and angles and extracting scenes based on these features, enabling highly accurate scene extraction even in different shooting environments. This enables efficient scene extraction according to the user's intentions and rapid analysis of game video.
[0006] "Video of a scene to be learned" is video used to extract features as a basis for analysis.
[0007] "Features" are important information for identifying a scene, such as specific movements, positioning, and ball trajectory present in the video.
[0008] "Game footage" refers to footage taken during an actual game, and is a continuous video that includes the progress of the game and the movements of the players.
[0009] "Extracting" means selecting and separating specific features or scenes from the video being analyzed.
[0010] A "user" is a person who uses this system and performs operations such as uploading videos, checking scenes, and providing feedback.
[0011] "Judge" means to check whether the scene presented to the user is what the user intended and to judge whether it is correct or not.
[0012] A "learning model" is a statistical, machine learning model that uses past data and feedback to improve the accuracy of analysis algorithms.
[0013] "Updating" is the act of adding new data or feedback to maintain or improve the performance of the system.
[0014] "Player movement" refers to the movement and movement patterns of players in the video.
[0015] "Positioning" refers to the positional information of players and the ball in the video, and the relationship between them.
[0016] "Ball trajectory" refers to the path and direction of travel of the ball in the video. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9]1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] System Overview
[0039] This invention is a system for efficiently analyzing amateur sports game footage. This system has the function of learning specific scenes and automatically extracting specific scenes from game footage based on the learning data. This allows users to quickly check the intended scenes without any effort, greatly improving the efficiency of the analysis process.
[0040] Program processing
[0041] Learning Phase
[0042] 1. The user saves the video of the scene they want to extract (e.g., a set play video during practice) on their device and uploads it to the server via a dedicated application.
[0043] 2. The server receives the uploaded video and begins analysis.
[0044] 3. The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory. These features are stored as tags in a database.
[0045] Match video analysis phase
[0046] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[0047] 2. The server receives the match video and begins video analysis.
[0048] 3. The server analyzes the game footage frame by frame and compares it with pre-trained features. During this analysis process, scenes that match the features are picked out.
[0049] 4. The selected scenes are organized into a list and temporarily saved on the server.
[0050] Verification Phase
[0051] 1. The user checks the scene clips extracted from the server through a dedicated application.
[0052] 2. The user plays each clip and judges whether it is the intended scene. This judgment is made by a dedicated application, with a "correct" or "incorrect" result.
[0053] 3. The user's judgment results are sent to the server and recorded in a database. This data is used to update the system's learning model and improve analysis accuracy in future analyses.
[0054] Specific examples
[0055] 1. The user takes a video of a "set play" practice (e.g., free kick practice) and saves the video on their device.
[0056] 2. The user uploads this video to the server using a dedicated application.
[0057] 3. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball trajectory and player positioning during a free kick are recorded as tags.
[0058] 4. The user takes video of the game and uploads it from their device to the server.
[0059] 5. The server receives the game footage and picks out scenes that match the learned features (e.g., free kick scenes).
[0060] 6. The server lists the scenes picked up and presents them to the user.
[0061] 7. The user uses a dedicated application to check the clips of each scene from the list and determine whether they are the intended free kick scenes.
[0062] 8. The user judges whether the result is "correct" or "incorrect," and this information is sent to the server. The server uses this data to update the learning model and improve the accuracy of future analyses.
[0063] In this way, the analysis of game footage is automated, significantly reducing the user's workload. Furthermore, the system performs incremental learning, which continuously improves analysis accuracy and enables more accurate scene extraction.
[0064] The processing flow will be explained below.
[0065] Learning Phase
[0066] Step 1:
[0067] The video of the scene the user wants to extract (e.g., a set play video during practice) is saved on the device.
[0068] Step 2:
[0069] The user launches the dedicated application and selects the video of the scene they want to extract.
[0070] Step 3:
[0071] The terminal uploads the selected video to the server.
[0072] Step 4:
[0073] The server prepares to analyze the received video.
[0074] Step 5:
[0075] The server analyzes the video frame by frame and extracts features such as players' movements, positioning, and ball trajectory.
[0076] Step 6:
[0077] The server stores the extracted features as tags in a database.
[0078] Match video analysis phase
[0079] Step 1:
[0080] The user saves the game video on the device.
[0081] Step 2:
[0082] The user launches the dedicated application and selects the game footage.
[0083] Step 3:
[0084] The terminal uploads the selected game video to the server.
[0085] Step 4:
[0086] The server prepares to analyze the received game footage.
[0087] Step 5:
[0088] The server analyzes the game footage frame by frame and compares it with pre-learned features.
[0089] Step 6:
[0090] The server picks out scenes that match the matched features.
[0091] Step 7:
[0092] The server organizes the picked scenes into a list and temporarily saves them.
[0093] Verification Phase
[0094] Step 1:
[0095] The user launches the dedicated application and checks the scene clips extracted from the server.
[0096] Step 2:
[0097] The user plays each clip and determines whether it is the intended scene.
[0098] Step 3:
[0099] The user selects "true" or "false" for each clip in a dedicated application.
[0100] Step 4:
[0101] The user's decision result is sent to the server.
[0102] Step 5:
[0103] The server records the received judgment result in a database.
[0104] Step 6:
[0105] The server uses the recorded data to update the learning model, improving the accuracy of analysis from the next time onwards.
[0106] Specific examples
[0107] Learning Phase
[0108] Step 1:
[0109] The user saves a "set play" practice video (e.g., free kick practice) on the device.
[0110] Step 2:
[0111] The user launches the dedicated application, selects a practice video, and uploads it to the server.
[0112] Step 3:
[0113] The device sends the video to the server.
[0114] Step 4:
[0115] The server receives the video and prepares it for analysis.
[0116] Step 5:
[0117] The server analyzes the video frame by frame and extracts features such as player movements, ball trajectory, and player positioning.
[0118] Step 6:
[0119] The server registers the extracted features in a database.
[0120] Match video analysis phase
[0121] Step 1:
[0122] The user saves the game video on the device.
[0123] Step 2:
[0124] The user launches the dedicated application, selects the game footage, and uploads it to the server.
[0125] Step 3:
[0126] The device sends the video to the server.
[0127] Step 4:
[0128] The server receives the match footage and prepares it for analysis.
[0129] Step 5:
[0130] The server analyzes the game footage frame by frame and compares it with the learned features.
[0131] Step 6:
[0132] The server picks out scenes that match the matched features (e.g., free kick scenes).
[0133] Step 7:
[0134] The server creates a list of the scenes it picks up and temporarily saves them.
[0135] Verification Phase
[0136] Step 1:
[0137] The user starts the dedicated application and checks the clip of the scene notified by the server.
[0138] Step 2:
[0139] The user plays each clip and checks whether it is the intended scene.
[0140] Step 3:
[0141] The user uses a dedicated application to make a judgment ("correct" or "incorrect") about each clip.
[0142] Step 4:
[0143] The user's decision result is sent to the server.
[0144] Step 5:
[0145] The server stores the received judgment results in a database.
[0146] Step 6:
[0147] The server uses the stored data to update the analysis model and improve accuracy from the next time onwards.
[0148] This automates the analysis of game footage, allowing users to efficiently review the scenes they intended. The system continues to learn, improving its analysis accuracy and providing an increasingly rich analysis experience.
[0149] Example 1
[0150] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0151] In conventional amateur sports game video analysis systems, extracting specific scenes is a manual process that requires a great deal of time and effort. Furthermore, the accuracy of scene analysis is low, making it difficult to identify the desired scene. Furthermore, there are problems with systems that cannot handle changes in the video due to shooting at different distances or angles.
[0152] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0153] In this invention, the server includes: means for receiving video of a scene to be learned and extracting features from the video; means for receiving game video and extracting scenes from the game video that match the features; means for presenting the extracted scenes to a user and allowing the user to determine whether the scene is a specific scene; means for recording the user's determination and updating the learning model to improve extraction accuracy in future runs; means for analyzing the video frame by frame and extracting player movements, positioning, and ball trajectory; means for storing the uploaded video in a storage system; and means for selecting whether the determination result is "correct" or "incorrect." This not only enables automatic extraction of specific scenes with high accuracy, but also enables adaptation to different shooting conditions. Users can quickly and effortlessly confirm the intended scene, achieving efficient analysis work.
[0154] "Means for receiving video of a scene to be learned and extracting features from the video" refers to a device or method in which a server takes in video data containing a specific scene provided by a user, and analyzes and extracts identifiable elements from the video, such as player movements, positioning, and ball trajectory.
[0155] The "means for receiving game footage and extracting scenes from the game footage that match the features" is a mechanism by which the server receives game footage provided by the user, compares it with features learned in advance, and automatically extracts matching parts.
[0156] The "means of presenting the extracted scene to the user and allowing the user to determine whether the scene is a specified scene" refers to a method in which the server displays the video clip extracted as the analysis result to the user through a user interface, allowing the user to check and determine whether the clip is the scene they want.
[0157] "Means for recording the user's judgment results and updating the learning model to improve extraction accuracy in the future" refers to the process in which the server stores the judgment results ("correct" or "incorrect") provided by the user in a database, and retrains and updates the machine learning model based on them to improve the accuracy of subsequent analysis.
[0158] "Means of analyzing video frame by frame to extract player movements, positioning, and ball trajectory" refers to a technology that divides video data into fixed time intervals and analyzes each frame in detail to extract important information such as the movements and positioning of players and the ball.
[0159] The "means for storing uploaded video in a storage system" refers to online storage or a database for temporarily or long-term storage of video data sent by a user to a server.
[0160] The "means for selecting whether the judgment result is 'correct' or 'incorrect'" is a user interface that allows the user to select 'correct' or 'incorrect' as to whether the content of a presented video clip matches a specified scene.
[0161] System Overview
[0162] This invention is a system for efficiently analyzing amateur sports game footage. This system has the function of learning specific scenes and automatically extracting specific scenes from game footage based on the learning data. This allows users to quickly check the intended scenes without any effort, greatly improving the efficiency of the analysis process.
[0163] Hardware and software used
[0164] The hardware used to implement this system includes general servers, terminals, and devices that provide user interfaces. The servers require a high-performance CPU and a large amount of storage, and it is recommended to use cloud services (e.g., AWS (registered trademark), Google (registered trademark) Cloud). Terminals can be smartphones, tablets, or PCs.
[0165] The software uses OpenCV for video analysis, TENSORFLOW (registered trademark) or PyTorch for machine learning, and MySQL (registered trademark) for database management. The dedicated application provides an interface for users to upload recorded video to the server and check the analysis results.
[0166] Program processing
[0167] Learning Phase
[0168] The user saves the video of the scene they want to extract (for example, a set play during practice) on their device. This video is then uploaded to the server via a dedicated application. The server analyzes the received video frame by frame, extracting features such as player movements, positioning, and ball trajectory, and stores these features as tags in a database.
[0169] Match video analysis phase
[0170] The user saves the game footage on their device and uploads it to the server via a dedicated application. The server receives the game footage and begins video analysis. The game footage is analyzed frame by frame and compared with pre-trained feature vectors. During this analysis process, scenes that match the feature vectors are picked out, organized into a list, and temporarily stored on the server.
[0171] Verification Phase
[0172] The user checks the scene clips extracted from the server through a dedicated application. Each clip is played and judged to be the intended scene. The judgement is made in the dedicated application as "correct" or "incorrect." The user's judgement is sent to the server and recorded in a database. This data is used to update the system's learning model and improve the accuracy of analysis from the next time onwards.
[0173] Specific examples
[0174] The user films a "set play" practice video (e.g., free kick practice) and saves the video on their device. The user then uploads this video to a server using a dedicated application. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball's trajectory and player positioning during a free kick are recorded as tags. The user films a game during the match and uploads the video from their device to the server. The server receives the game video and selects scenes (e.g., free kick scenes) that match the learned features. The server then lists the selected scenes and presents them to the user. The user then uses a dedicated application to review each scene clip from the list and determine whether it is the intended free kick scene. The user then judges whether it is "correct" or "incorrect," and this information is sent to the server. The server uses this data to update the learning model, improving the accuracy of future analysis.
[0175] Prompt Sentence Examples
[0176] "Please explain how to film amateur soccer free kick practice scenes, save them on a device, and upload the footage to a server using a dedicated application. Also, please explain the steps for how the server analyzes and learns from the uploaded footage."
[0177] In this way, the analysis of game footage is automated, significantly reducing the user's workload. Furthermore, the system performs incremental learning, which continuously improves analysis accuracy and enables more accurate scene extraction.
[0178] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0179] Step 1:
[0180] The video of the scene that the user wants to extract (e.g., a set play video during practice) is recorded and saved on the device. The input is a video file containing a specific scene created by the user, and the output is a video file saved on the device. The user uses a smartphone or camera to record the set practice scene and saves the data on the device.
[0181] Step 2:
[0182] The user launches the dedicated application and selects the recorded video file within the application. Then, they click the "Upload" button in the application to upload the video file to the server. The input is the video file saved on the device, and the output is the video data uploaded to the server.
[0183] Step 3:
[0184] The server receives the uploaded video file. It stores the received video file in a storage system (e.g., AWS S3). The input is the uploaded video file, and the output is the stored video data.
[0185] Step 4:
[0186] The server analyzes the video frame by frame. Video processing software such as OpenCV is used for video analysis. Through the analysis, features such as player movements, positioning, and ball trajectory are extracted and stored as tags in a database. The input is the stored video data, and the output is tag data of the extracted features.
[0187] Step 5:
[0188] The user saves the game video on the device. The entire game video is recorded and the data is saved on the device. The input is the game video, and the output is the game video saved on the device. The user also records the game video and saves it on the device.
[0189] Step 6:
[0190] The user uploads game footage saved on their device to the server using a dedicated application. The input is the game footage file saved on the device, and the output is the game footage data uploaded to the server.
[0191] Step 7:
[0192] The server receives and stores the game footage. It analyzes the received footage frame by frame and compares it with pre-trained features. Computer vision technology and machine learning models (e.g., TensorFlow) are used for the matching process. The input is the stored game footage data, and the output is the extraction of scenes that match the features.
[0193] Step 8:
[0194] The server compiles a list of matched scenes and temporarily stores it. The input is the frame data of the matched scenes, and the output is a scene list. These scenes are organized for presentation to the user.
[0195] Step 9:
[0196] The user launches a dedicated application and checks the scene list provided by the server. The user selects the "Check Scene" option within the application and plays a clip of the presented scene. The input is the list provided by the server, and the output is a clip of the scene selected by the user.
[0197] Step 10:
[0198] The user plays each clip and judges whether it is the intended scene as "correct" or "incorrect." This judgment is made on the interface of a dedicated application. The input is the played clip video, and the output is the judgment result, "correct" or "incorrect."
[0199] Step 11:
[0200] The user's judgment results are sent to the server via a dedicated application. The server stores the judgment results in a database and uses this data to retrain and update the machine learning model. The input is the user's judgment results, and the output is the updated learning model. This improves the accuracy of analysis from the next time onwards.
[0201] (Application example 1)
[0202] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0203] This invention relates to a system that efficiently analyzes the manufacturing processes of automated robots and workers used in factories, automatically extracts and evaluates specific work scenes, and improves manufacturing efficiency and detects anomalies. In current manufacturing processes, it is difficult to quickly detect and analyze inefficient operations or abnormalities, which can result in adverse effects on product quality and productivity. To solve this problem, technology is needed that can automatically monitor and analyze manufacturing processes and extract specific work scenes.
[0204] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0205] In this invention, the server includes means for receiving video of a work scene to be learned and extracting features from the video, means for receiving manufacturing process video and extracting scenes from the manufacturing process video that match the features, means for presenting the extracted work scene to a user and allowing the user to determine whether the scene is a predetermined work scene, means for recording the user's determination result and updating the learning model to improve extraction accuracy from the next time onwards, means for analyzing scenes suspected of abnormalities or reduced efficiency and organizing them into a list, and means for generating improvement suggestions based on the analysis results, thereby enabling efficiency improvement in the manufacturing process and early detection of abnormalities.
[0206] "Video of a work scene to be learned" is video data that records the process of a specific work being done in a factory.
[0207] "Features" are data extracted from video data, such as the worker's movements, frequency of tool use, and work time.
[0208] "Manufacturing process video" is video data that records the actual manufacturing process.
[0209] A "matching scene" is a part of the manufacturing process video that has features similar to the learned features.
[0210] "User" refers to a person in a factory who uses the system to analyze the manufacturing process.
[0211] A "predetermined work scene" is a video scene showing the process of a specific work that has been determined in advance.
[0212] The "determination result" is the result of the user's determination as to whether or not the scene is a predetermined work scene.
[0213] A "learning model" is a data processing algorithm that learns features based on video data and improves analysis accuracy.
[0214] "Scenes where abnormalities or reduced efficiency are suspected" are scenes where abnormal operations or reduced work efficiency are observed compared to the normal manufacturing process.
[0215] A "list" is data that organizes and lists situations where abnormalities or reduced efficiency are suspected.
[0216] "Improvement proposals" are specific action plans proposed based on the analysis results to improve the efficiency of manufacturing processes and correct abnormalities.
[0217] The system for implementing this invention starts by recording the manufacturing process of automated robots and workers used in a factory with a camera and uploading the video data to a server. The server then uses dedicated software for analyzing the video data (e.g., OpenCV or TensorFlow) to extract features such as the worker's movements, frequency of tool use, and work time for each frame and stores them in a database.
[0218] Next, actual manufacturing process footage is similarly uploaded to the server and compared with the pre-trained feature values. At this time, the server picks out scenes that are suspected of being abnormal or inefficient and organizes them into a list. The user checks this list using a dedicated device (tablet or smartphone) and determines whether each scene is a specified work scene. The results of this determination are sent to the server, and the learning model is updated, improving extraction accuracy from the next time onwards.
[0219] Furthermore, the server generates improvement proposals based on the analysis results and provides them to users. This enables the efficiency of the manufacturing process and the early detection of abnormalities. In addition, the system has the advantage of being able to continuously improve its analysis accuracy because it learns sequentially.
[0220] Hardware used
[0221] Camera (installed inside the factory)
[0222] Factory Robots
[0223] Dedicated tablet / smartphone
[0224] Servers (e.g., high-performance computing resources on Amazon Web Services (AWS) or Google Cloud Platform (GCP))
[0225] Software used
[0226] Analysis software (e.g., OpenCV (image processing library), TensorFlow (machine learning library))
[0227] Database management system (e.g. MySQL)
[0228] Specific examples
[0229] For example, to analyze a video of a scene in an assembly room in a factory where a robot arm is performing a specific movement and detect a specific movement pattern, the following prompt sentence would be used.
[0230] Prompt Sentence Examples
[0231] "During factory assembly work, identify instances where the robot arm's movement slows down or stops, and compare it with manual work."
[0232] Based on this prompt, the generative AI model analyzes the video data from within the factory, distinguishes between efficient and abnormal operations, and outputs improvement suggestions, allowing users to quickly identify problems in the manufacturing process and take measures.
[0233] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0234] Step 1:
[0235] Users use a dedicated device (tablet or smartphone) to record video of specific tasks within the factory and save it on the device.
[0236] Input: Video data of specific work being done in a factory
[0237] Output: Work video data saved on the device
[0238] Step 2:
[0239] The user uploads the work video data to the server through the terminal.
[0240] Input: Work video data saved on the device
[0241] Output: Work video data uploaded to the server
[0242] Step 3:
[0243] The server receives the uploaded work video data and extracts features using analysis software (e.g., OpenCV or TensorFlow).
[0244] Input: Work video data uploaded to the server
[0245] Output: Extracted features (worker movements, tool usage frequency, work time, etc.)
[0246] Step 4:
[0247] The server stores the extracted features as tags in a database.
[0248] Input: Extracted features
[0249] Output: Tag information stored in the database
[0250] Step 5:
[0251] The user uploads the actual manufacturing process video to the server via the terminal.
[0252] Input: Actual manufacturing process video data
[0253] Output: Manufacturing process video data uploaded to the server
[0254] Step 6:
[0255] The server analyzes the received manufacturing process video data and compares it with the aforementioned features.
[0256] Input: Manufacturing process video data uploaded to the server and tag information stored in the database
[0257] Output: Matching result (scenes that match the features)
[0258] Step 7:
[0259] Identify situations where the server is suspected to be abnormal or inefficient and organize them into a list.
[0260] Input: Matching result
[0261] Output: A list of suspected anomalies and inefficiencies
[0262] Step 8:
[0263] The user checks the list using a dedicated terminal and determines whether each scene is a predetermined work scene.
[0264] Input: List of suspected anomalies or inefficiencies
[0265] Output: Judgment result (whether it is a given work scene or not)
[0266] Step 9:
[0267] The judgment result is sent to the server, and the server updates the learning model.
[0268] Input: Judgment result
[0269] Output: Updated training model
[0270] Step 10:
[0271] The server generates improvement suggestions based on the analysis results and provides them to the user.
[0272] Input: Updated training model and analysis results
[0273] Output: Improvement suggestions
[0274] In this way, specific operations and data processing are carried out at each step, resulting in more efficient manufacturing processes and early detection of abnormalities.
[0275] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0276] System Overview
[0277] This invention is a system for learning and extracting specific scenes for the purpose of efficiently analyzing amateur sports game footage. This system has the ability to learn specific patterns and automatically extract specific scenes from game footage. Furthermore, by combining it with an emotion engine that recognizes user emotions and adjusts the analysis results based on those emotions, it is possible to improve user satisfaction.
[0278] Program processing
[0279] Learning Phase
[0280] 1. The user saves the video of the scene they want to extract (e.g., a set play video during practice) on their device and uploads it to the server via a dedicated application.
[0281] 2. The server receives the uploaded video and begins analysis.
[0282] 3. The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory. These features are stored as tags in a database.
[0283] Match video analysis phase
[0284] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[0285] 2. The server receives the match video and begins video analysis.
[0286] 3. The server analyzes the game footage frame by frame and compares it with pre-trained features. During this analysis process, scenes that match the features are picked out.
[0287] 4. The selected scenes are organized into a list and temporarily saved on the server.
[0288] Emotion Recognition Phase
[0289] 1. When a user uses a dedicated application to check the extracted scene clip, the device analyzes the user's facial expressions, voice, and gestures using an emotion engine.
[0290] 2. The emotion engine recognizes the user's emotions and sends that information to the server.
[0291] 3. The server receives the emotion engine data and dynamically adjusts the content and display method of the presented scene based on the user's emotions.
[0292] 4. The server records changes in the user's emotions and uses this information to improve the way scenes are presented and the accuracy of the analysis content in future sessions.
[0293] Verification Phase
[0294] 1. The user plays each clip and judges whether it is the intended scene. This judgment is made by a dedicated application, with a "correct" or "incorrect" result.
[0295] 2. The user's judgment results are sent to the server and recorded in a database. This data is used to update the system's learning model and improve the accuracy of analysis in future.
[0296] Specific examples
[0297] 1. The user takes a video of a "set play" practice (e.g., free kick practice) and saves the video on their device.
[0298] 2. The user uploads this video to the server using a dedicated application.
[0299] 3. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball trajectory and player positioning during a free kick are recorded as tags.
[0300] 4. The user takes video of the game and uploads it from their device to the server.
[0301] 5. The server receives the game footage and picks out scenes that match the learned features (e.g., free kick scenes).
[0302] 6. The server lists the scenes picked up and presents them to the user.
[0303] 7. When the user uses a dedicated application to check clips for each scene from the list, the device analyzes the user's facial expressions, voice, and gestures, and the emotion engine recognizes the user's emotions.
[0304] 8. The server dynamically adjusts how the clip is presented based on the user's emotions, for example, adjusting the playback speed of the clip if a positive emotion is detected, or revisiting a section of the clip if a negative emotion is detected.
[0305] 9. The user marks each clip as "true" or "false," and the result is sent to the server.
[0306] 10. The server records the judgment data and emotion data from the user, updates the system's learning model, and improves the accuracy of analysis from the next time onwards.
[0307] In this way, game video analysis is automated and dynamic scene adjustments based on the user's emotions are made possible, providing a more intuitive and satisfying video analysis experience. The system continues to learn and improves its analysis accuracy, resulting in an increasingly rich analysis experience.
[0308] The processing flow will be explained below.
[0309] Learning Phase
[0310] Step 1:
[0311] The video of the scene the user wants to extract (e.g., a set play video during practice) is saved on the device.
[0312] Step 2:
[0313] The user launches the dedicated application and selects the video of the scene they want to extract.
[0314] Step 3:
[0315] The terminal uploads the selected video to the server.
[0316] Step 4:
[0317] The server receives the uploaded video and begins analysis.
[0318] Step 5:
[0319] The server analyzes the video frame by frame and extracts features such as players' movements, positioning, and ball trajectory.
[0320] Step 6:
[0321] The server stores the extracted features as tags in a database.
[0322] Match video analysis phase
[0323] Step 1:
[0324] The user saves the game video on the device.
[0325] Step 2:
[0326] The user launches the dedicated application and selects the game footage.
[0327] Step 3:
[0328] The terminal uploads the selected game video to the server.
[0329] Step 4:
[0330] The server prepares to analyze the received game footage.
[0331] Step 5:
[0332] The server analyzes the game footage frame by frame and compares it with pre-learned features.
[0333] Step 6:
[0334] The server picks out scenes that match the matched features.
[0335] Step 7:
[0336] The server organizes the picked scenes into a list and temporarily saves them.
[0337] Emotion Recognition Phase
[0338] Step 1:
[0339] The user launches the dedicated application and checks the scene clips extracted from the server.
[0340] Step 2:
[0341] The device uses an emotion engine to analyze the user's emotional information, such as facial expressions, voice, and gestures.
[0342] Step 3:
[0343] The emotion engine recognizes the user's emotions and sends the information to the server.
[0344] Step 4:
[0345] The server receives the emotion engine data and dynamically adjusts the presentation method and content of the extracted scene based on the user's emotion.
[0346] Step 5:
[0347] The server records changes in the user's emotions and uses this information to improve the accuracy of scene presentations and analysis content in future sessions.
[0348] Verification Phase
[0349] Step 1:
[0350] The user plays each clip and determines whether it is the intended scene.
[0351] Step 2:
[0352] The user selects "true" or "false" for each clip in a dedicated application.
[0353] Step 3:
[0354] The user's decision result is sent to the server.
[0355] Step 4:
[0356] The server stores the judgment results and emotion recognition data in a database.
[0357] Step 5:
[0358] The server uses the stored data to update the learning model and improve the accuracy of analysis from the next time onwards.
[0359] Specific examples
[0360] Learning Phase
[0361] Step 1:
[0362] The user saves a "set play" practice video (e.g., free kick practice) on the device.
[0363] Step 2:
[0364] The user launches the dedicated application, selects a practice video, and uploads it to the server.
[0365] Step 3:
[0366] The device sends the video to the server.
[0367] Step 4:
[0368] The server receives the video and prepares it for analysis.
[0369] Step 5:
[0370] The server analyzes the video frame by frame and extracts features such as player movements, ball trajectory, and player positioning.
[0371] Step 6:
[0372] The server registers the extracted features in a database.
[0373] Match video analysis phase
[0374] Step 1:
[0375] The user saves the game video on the device.
[0376] Step 2:
[0377] The user launches the dedicated application, selects the game footage, and uploads it to the server.
[0378] Step 3:
[0379] The device sends the video to the server.
[0380] Step 4:
[0381] The server receives the match footage and prepares it for analysis.
[0382] Step 5:
[0383] The server analyzes the game footage frame by frame and compares it with the learned features.
[0384] Step 6:
[0385] The server picks out scenes that match the matched features (e.g., free kick scenes).
[0386] Step 7:
[0387] The server creates a list of the scenes it picks up and temporarily saves them.
[0388] Emotion Recognition Phase
[0389] Step 1:
[0390] The user starts the dedicated application and checks the clips of the scenes listed from the server.
[0391] Step 2:
[0392] The device analyzes the user's facial expressions, voice, gestures, etc. using an emotion engine.
[0393] Step 3:
[0394] The emotion engine recognizes the user's emotion and transmits the emotion information to the server.
[0395] Step 4:
[0396] The server receives the emotion engine data and dynamically adjusts how the clip is displayed based on the user's emotion.
[0397] Step 5:
[0398] The server records the user's emotional fluctuations and uses this information to improve the way scenes are presented and the accuracy of the analysis content from next time onwards.
[0399] Verification Phase
[0400] Step 1:
[0401] The user plays each clip and checks whether it is the intended scene.
[0402] Step 2:
[0403] The user selects "true" or "false" for each clip and sends the result to the server.
[0404] Step 3:
[0405] The server stores the judgment results and emotion recognition data in a database.
[0406] Step 4:
[0407] The server uses the stored data to update the learning model and improve analysis accuracy.
[0408] This series of processes automates the analysis of match footage and enables dynamic scene adjustments based on the user's emotions, providing a more intuitive and satisfying video analysis experience. The system continues to learn and improves its analysis accuracy, resulting in an increasingly rich analysis experience.
[0409] Example 2
[0410] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0411] When analyzing amateur sports game footage, it is difficult to efficiently and accurately extract specific scenes. Furthermore, to improve user satisfaction with the information obtained from video analysis, it is necessary to dynamically adjust the video based on the user's emotions. However, current systems lack the functionality to recognize the user's emotions and dynamically adjust the video presentation based on those emotions. As a result, users may not obtain the information they expect, potentially reducing the effectiveness of video analysis. A system that can solve these issues is needed.
[0412] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0413] In this invention, the server includes means for receiving video of scenes to be learned and extracting features from the video, means for receiving game video and extracting scenes from the game video that match the features, means for presenting the extracted scenes to a user and allowing the user to determine whether the scenes are predetermined scenes, means for recording the user's determination result and updating the learning model to improve extraction accuracy in the future, and means for recognizing the user's emotions and dynamically adjusting the content and presentation method of the extracted scenes based on the recognized emotions. This enables efficient and accurate scene extraction from game video and dynamic adjustment of the video in accordance with the user's emotions, thereby achieving video analysis with high user satisfaction.
[0414] Below are definitions of important words.
[0415] "Video of a scene to be learned" refers to video data that includes a specific scene that the user wants to extract.
[0416] "Features" refers to information extracted from video data that is necessary to identify specific scenes, such as player movements, positioning, and ball trajectory.
[0417] "Game footage" refers to video data that records the progress of an actual sports game.
[0418] An "emotion engine" refers to software or hardware that analyzes a user's facial expressions, voice, and gestures to recognize emotions.
[0419] A "learning model" refers to the algorithms and data structures that allow a system to continuously learn based on past data.
[0420] "Server" refers to a computer system that processes and stores video data and analysis data uploaded by users and performs various analyses.
[0421] "Terminal" refers to the device (e.g., smartphone, tablet, PC, etc.) that a user uses to store video and operate a dedicated application.
[0422] "Tag" refers to a label or metadata that is added to identify a specific feature in a video.
[0423] "Extraction" refers to extracting features or scenes that match predetermined conditions from video data.
[0424] "Dynamic adjustment" refers to changing the way the video is displayed in real time based on the user's emotions and other variables.
[0425] This allows the technical scope of the invention to be more clearly defined.
[0426] This invention is a system for efficiently analyzing amateur sports game footage, and has the ability to learn and extract specific scenes. This system learns specific patterns and automatically extracts specific scenes from game footage. Furthermore, by combining it with an emotion engine that recognizes user emotions and adjusts the analysis results based on those emotions, it is possible to improve user satisfaction.
[0427] The main hardware and software components for implementing this system are as follows:
[0428] 1. Device: A device used by a user to store video and operate dedicated applications. Examples include smartphones, tablets, and PCs.
[0429] 2. Server: A computer system that processes and stores video data and analysis data uploaded by users and performs various analyses. It is responsible for video analysis, database management, and execution of learning models.
[0430] 3. Emotion engine: Software or hardware that analyzes the user's facial expressions, voice, and gestures to recognize emotions.
[0431] The specific steps for implementing the system are as follows:
[0432] Learning Phase
[0433] 1. The user takes a video of the scene they want to extract and saves it on their device.
[0434] 2. The user uploads the video to the server through a dedicated application.
[0435] 3. The server receives the video and analyzes it frame by frame to extract features such as player movements, positioning, and ball trajectory.
[0436] Match video analysis phase
[0437] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[0438] 2. The server receives the game footage, analyzes each frame, and compares it with pre-trained features.
[0439] 3. The server picks up matching scenes, organizes them in a list format, and temporarily saves them.
[0440] Emotion Recognition Phase
[0441] 1. The user uses a dedicated application to view clips of the scenes selected here.
[0442] 2. The device sends the user's facial expressions, voice, and gestures to the emotion engine.
[0443] 3. The emotion engine generates emotion data and sends it to the server.
[0444] 4. The server dynamically adjusts the content and presentation of the clip based on the emotional data.
[0445] Verification Phase
[0446] 1. The user judges each clip as "true" or "false."
[0447] 2. The user's judgment result is sent to the server and recorded in the database.
[0448] 3. The server uses the judgment results to update the learning model and improve the accuracy of analysis from the next time onwards.
[0449] For example, if a user films and uploads a free-kick practice video, the server analyzes the ball's trajectory and player positions to select free-kick scenes from game footage. Then, when the user plays back the video, the system can dynamically adjust the playback speed and content of the clip based on the user's emotions.
[0450] Prompt Sentence Examples
[0451] 1. "Please explain the scenario where a user uploads a video of themselves practicing free kicks."
[0452] 2. "Please explain how to analyze frames from game footage."
[0453] 3. "Please explain how the emotion engine recognizes the user's emotions and how the server uses them."
[0454] This system enables efficient and accurate extraction of specific scenes, and also realizes dynamic adjustment of the video according to the user's emotions, thereby providing video analysis with high user satisfaction.
[0455] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0456] Learning Phase
[0457] Step 1:
[0458] The user takes a video of the scene they want to extract and saves it on their device.
[0459] Input: Video files captured by the camera.
[0460] Output: Video data saved on the device.
[0461] Step 2:
[0462] The user opens the dedicated application, selects the saved video file, and uploads it to the server.
[0463] Input: User specified video file.
[0464] Output: Video data received by the server.
[0465] Step 3:
[0466] The server starts a program to analyze the received video file.
[0467] Input: Uploaded video file.
[0468] Output: Video data divided into frames.
[0469] Step 4:
[0470] The server extracts features such as player movements, positioning, and ball trajectory for each frame.
[0471] Input: Video data divided into frames.
[0472] Output: Extracted feature data (player movements and positions, ball trajectory, etc.).
[0473] Step 5:
[0474] The server stores the extracted features as tags in a database.
[0475] Input: Extracted feature data.
[0476] Output: Tag information stored in the database.
[0477] Match video analysis phase
[0478] Step 1:
[0479] Users save game footage on their devices and upload it to the server via a dedicated application.
[0480] Input: User-specified game video file.
[0481] Output: Match video data received by the server.
[0482] Step 2:
[0483] The server receives the game footage and starts a program that analyzes it frame by frame.
[0484] Input: Uploaded match video file.
[0485] Output: Match video data divided into frames.
[0486] Step 3:
[0487] The server compares the features with those learned in advance and picks out scenes with matching features.
[0488] Input: Game video data divided into frames and pre-trained feature data.
[0489] Output: Frame number and time information of the matching scenes.
[0490] Step 4:
[0491] The server organizes the picked-up scenes in list format and temporarily saves them.
[0492] Input: Frame number and time information of the matching scene.
[0493] Output: List of scene information and its temporary storage.
[0494] Emotion Recognition Phase
[0495] Step 1:
[0496] The user can use a dedicated application to check clips of the scenes picked up here.
[0497] Input: Listed scene information.
[0498] Output: The clip of the scene being played.
[0499] Step 2:
[0500] The device transmits the user's facial expressions, voice, and gestures to the emotion engine.
[0501] Input: User facial, voice, and gesture data.
[0502] Output: Data sent to the emotion engine.
[0503] Step 3:
[0504] The emotion engine generates emotion data and sends it to the server.
[0505] Input: User facial, voice, and gesture data.
[0506] Output: The generated emotion data.
[0507] Step 4:
[0508] The server dynamically adjusts the content and presentation method of the clip based on the emotional data received.
[0509] Input: Emotion data and playing clip information.
[0510] Output: The playback speed and display of the adjusted clip.
[0511] Verification Phase
[0512] Step 1:
[0513] The user marks each clip as "true" or "false."
[0514] Input: The clip that was played.
[0515] Output: User's decision result.
[0516] Step 2:
[0517] The user's judgment result is sent to the server and recorded in a database.
[0518] Input: User's decision result.
[0519] Output: Verdict data recorded in a database.
[0520] Step 3:
[0521] The server uses the judgment results to update the learning model, improving the accuracy of analysis from the next time onwards.
[0522] Input: Recorded adjudication data.
[0523] Output: Updated training model.
[0524] (Application example 2)
[0525] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0526] When analyzing video of sports games, advertisements, etc., there is a need to efficiently extract specific scenes and dynamically adjust the display content based on the user's emotions. However, conventional technologies have had difficulty recognizing the user's emotions in real time and optimally displaying scenes and delivering advertisements in accordance with those emotions. The present invention aims to solve this problem and provide more intuitive and satisfying video analysis and advertisement delivery.
[0527] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0528] In this invention, the server includes means for receiving video of scenes to be learned and extracting features from the video, means for receiving game video and extracting scenes from the game video that match the features, and means for recognizing the user's emotions while watching and dynamically adjusting the content and display method of the extracted scenes based on the emotions, thereby enabling scene display and advertisement delivery in accordance with the user's emotions.
[0529] "Video of the scene to be learned" refers to video that the system uses to extract features in advance.
[0530] "Features" are data such as player movements, positioning, and ball trajectory that the system extracts from the video.
[0531] "Game footage" refers to footage of an actual sport or event being played.
[0532] "User" means any person or entity that uses the System.
[0533] "Means for adjusting based on emotions" is a mechanism that analyzes the user's emotions and dynamically changes the content and method of displaying the video based on the results.
[0534] "Dynamic adjustment" refers to changing the content and display method of the video in real time in response to changes in the user's emotions.
[0535] A "server" refers to a computer or network system that performs the main processing of a system.
[0536] System Overview
[0537] This invention is a system for efficiently analyzing sports game footage and advertising videos. It is characterized by recognizing the user's emotions while watching and dynamically adjusting the content and display method of the video based on those emotions. This system extracts features from the video of the scene being studied and analyzes the game footage and advertising videos in real time to provide the user with an optimal experience.
[0538] Program processing
[0539] Learning Phase
[0540] The server receives the video of the scene the user wants to extract (the learning video) and extracts features from the video. Specifically, it analyzes the video frame by frame and uses a generative AI model to extract data such as player movements, positioning, and ball trajectory. These features are then stored as tags in a database.
[0541] Match video analysis phase
[0542] Users save game footage or advertising footage on their devices and upload it to the server via a dedicated application. The server analyzes the received footage and compares it with pre-trained features to automatically select scenes that match. These scenes are organized into a list and temporarily saved.
[0543] Emotion Recognition Phase
[0544] When a user reviews an extracted scene using a dedicated application, the device analyzes the user's facial expressions, voice, and gestures using an emotion engine. The server receives the data sent from the emotion engine and dynamically adjusts how the scene is displayed based on the user's emotions. If a positive emotion is recognized, the playback speed of the clip is adjusted, and if a negative emotion is recognized, the clip is reviewed.
[0545] Hardware and software used
[0546] Hardware: High-performance servers, user devices (smartphones, tablets, PCs, etc.)
[0547] Software: Generative AI models (Keras, TensorFlow, OpenCV), emotion engine, database management system
[0548] Specific examples
[0549] For example, the present invention can be implemented as a system for delivering advertisements based on user emotions in the advertising industry. The server analyzes the advertisement video being viewed by the user in real time, recognizes the user's emotions while viewing, and dynamically selects the next advertisement to be delivered based on the emotions.
[0550] Examples of prompts:
[0551] "Please build a system that analyzes the advertisements that users watch in real time, recognizes their emotions from their facial expressions while watching, and dynamically selects the next advertisement to be delivered."
[0552] This makes it possible to deliver optimal video and advertisements according to the user's emotions, improving the viewing experience and maximizing the effectiveness of advertising.
[0553] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0554] Step 1:
[0555] Input: The user saves the video of the scene to be learned on the device.
[0556] How it works: The user uses a dedicated application to upload the video to be studied to the server.
[0557] Output: The learning video is saved on the server.
[0558] Step 2:
[0559] Input: The training video received by the server.
[0560] How it works: The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory, using generative AI models (Keras, TensorFlow).
[0561] Output: The extracted features are stored as tags in a database.
[0562] Step 3:
[0563] Input: The user saves game footage or advertising footage to the device.
[0564] How it works: Users use a dedicated application to upload game footage and advertising footage to the server.
[0565] Output: Match footage and advertising footage are saved on the server.
[0566] Step 4:
[0567] Input: Game footage and advertising footage received by the server, as well as pre-trained features.
[0568] How it works: The server analyzes the received video frame by frame and compares it with pre-trained features to pick out matching scenes. It then uses a generative AI model to confirm feature matches.
[0569] Output: Matching scenes are organized into a list and temporarily stored on the server.
[0570] Step 5:
[0571] Input: The user checks the extracted scene clips in a dedicated application.
[0572] How it works: While the user plays the clip, the device uses the camera and microphone to input the user's facial expressions, voice, and gestures into the emotion engine in real time.
[0573] Output: The emotion engine extracts the user's emotion data and sends it to the server.
[0574] Step 6:
[0575] Input: Emotion data received by the server, and a list of extracted scenes.
[0576] How it works: The server dynamically adjusts how the clip is presented based on the user's emotional data, for example adjusting the clip's playback speed if a positive emotion is detected, or revisiting a section of the clip if a negative emotion is detected.
[0577] Output: The adjusted scene is displayed to the user in real time.
[0578] Step 7:
[0579] Input: The user judges each clip as "true" or "false."
[0580] How it works: The user uses a dedicated application to send the results of their assessment of the clip to the server.
[0581] Output: The server records the user's judgment results in a database and updates the learning model to improve analysis accuracy in future.
[0582] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0583] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0584] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0585] [Second embodiment]
[0586] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0587] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0588] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0589] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0590] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0591] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0592] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0593] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0594] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0595] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0596] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0597] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0598] System Overview
[0599] This invention is a system for efficiently analyzing amateur sports game footage. This system has the function of learning specific scenes and automatically extracting specific scenes from game footage based on the learning data. This allows users to quickly check the intended scenes without any effort, greatly improving the efficiency of the analysis process.
[0600] Program processing
[0601] Learning Phase
[0602] 1. The user saves the video of the scene they want to extract (e.g., a set play video during practice) on their device and uploads it to the server via a dedicated application.
[0603] 2. The server receives the uploaded video and begins analysis.
[0604] 3. The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory. These features are stored as tags in a database.
[0605] Match video analysis phase
[0606] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[0607] 2. The server receives the match video and begins video analysis.
[0608] 3. The server analyzes the game footage frame by frame and compares it with pre-trained features. During this analysis process, scenes that match the features are picked out.
[0609] 4. The selected scenes are organized into a list and temporarily saved on the server.
[0610] Verification Phase
[0611] 1. The user checks the scene clips extracted from the server through a dedicated application.
[0612] 2. The user plays each clip and judges whether it is the intended scene. This judgment is made by a dedicated application, with a "correct" or "incorrect" result.
[0613] 3. The user's judgment results are sent to the server and recorded in a database. This data is used to update the system's learning model and improve analysis accuracy in future analyses.
[0614] Specific examples
[0615] 1. The user takes a video of a "set play" practice (e.g., free kick practice) and saves the video on their device.
[0616] 2. The user uploads this video to the server using a dedicated application.
[0617] 3. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball trajectory and player positioning during a free kick are recorded as tags.
[0618] 4. The user takes video of the game and uploads it from their device to the server.
[0619] 5. The server receives the game footage and picks out scenes that match the learned features (e.g., free kick scenes).
[0620] 6. The server lists the scenes picked up and presents them to the user.
[0621] 7. The user uses a dedicated application to check the clips of each scene from the list and determine whether they are the intended free kick scenes.
[0622] 8. The user judges whether the result is "correct" or "incorrect," and this information is sent to the server. The server uses this data to update the learning model and improve the accuracy of future analyses.
[0623] In this way, the analysis of game footage is automated, significantly reducing the user's workload. Furthermore, the system performs incremental learning, which continuously improves analysis accuracy and enables more accurate scene extraction.
[0624] The processing flow will be explained below.
[0625] Learning Phase
[0626] Step 1:
[0627] The video of the scene the user wants to extract (e.g., a set play video during practice) is saved on the device.
[0628] Step 2:
[0629] The user launches the dedicated application and selects the video of the scene they want to extract.
[0630] Step 3:
[0631] The terminal uploads the selected video to the server.
[0632] Step 4:
[0633] The server prepares to analyze the received video.
[0634] Step 5:
[0635] The server analyzes the video frame by frame and extracts features such as players' movements, positioning, and ball trajectory.
[0636] Step 6:
[0637] The server stores the extracted features as tags in a database.
[0638] Match video analysis phase
[0639] Step 1:
[0640] The user saves the game video on the device.
[0641] Step 2:
[0642] The user launches the dedicated application and selects the game footage.
[0643] Step 3:
[0644] The terminal uploads the selected game video to the server.
[0645] Step 4:
[0646] The server prepares to analyze the received game footage.
[0647] Step 5:
[0648] The server analyzes the game footage frame by frame and compares it with pre-learned features.
[0649] Step 6:
[0650] The server picks out scenes that match the matched features.
[0651] Step 7:
[0652] The server organizes the picked scenes into a list and temporarily saves them.
[0653] Verification Phase
[0654] Step 1:
[0655] The user launches the dedicated application and checks the scene clips extracted from the server.
[0656] Step 2:
[0657] The user plays each clip and determines whether it is the intended scene.
[0658] Step 3:
[0659] The user selects "true" or "false" for each clip in a dedicated application.
[0660] Step 4:
[0661] The user's decision result is sent to the server.
[0662] Step 5:
[0663] The server records the received judgment result in a database.
[0664] Step 6:
[0665] The server uses the recorded data to update the learning model, improving the accuracy of analysis from the next time onwards.
[0666] Specific examples
[0667] Learning Phase
[0668] Step 1:
[0669] The user saves a "set play" practice video (e.g., free kick practice) on the device.
[0670] Step 2:
[0671] The user launches the dedicated application, selects a practice video, and uploads it to the server.
[0672] Step 3:
[0673] The device sends the video to the server.
[0674] Step 4:
[0675] The server receives the video and prepares it for analysis.
[0676] Step 5:
[0677] The server analyzes the video frame by frame and extracts features such as player movements, ball trajectory, and player positioning.
[0678] Step 6:
[0679] The server registers the extracted features in a database.
[0680] Match video analysis phase
[0681] Step 1:
[0682] The user saves the game video on the device.
[0683] Step 2:
[0684] The user launches the dedicated application, selects the game footage, and uploads it to the server.
[0685] Step 3:
[0686] The device sends the video to the server.
[0687] Step 4:
[0688] The server receives the match footage and prepares it for analysis.
[0689] Step 5:
[0690] The server analyzes the game footage frame by frame and compares it with the learned features.
[0691] Step 6:
[0692] The server picks out scenes that match the matched features (e.g., free kick scenes).
[0693] Step 7:
[0694] The server creates a list of the scenes it picks up and temporarily saves them.
[0695] Verification Phase
[0696] Step 1:
[0697] The user starts the dedicated application and checks the clip of the scene notified by the server.
[0698] Step 2:
[0699] The user plays each clip and checks whether it is the intended scene.
[0700] Step 3:
[0701] The user uses a dedicated application to make a judgment ("correct" or "incorrect") about each clip.
[0702] Step 4:
[0703] The user's decision result is sent to the server.
[0704] Step 5:
[0705] The server stores the received judgment results in a database.
[0706] Step 6:
[0707] The server uses the stored data to update the analysis model and improve accuracy from the next time onwards.
[0708] This automates the analysis of game footage, allowing users to efficiently review the scenes they intended. The system continues to learn, improving its analysis accuracy and providing an increasingly rich analysis experience.
[0709] Example 1
[0710] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0711] In conventional amateur sports game video analysis systems, extracting specific scenes is a manual process that requires a great deal of time and effort. Furthermore, the accuracy of scene analysis is low, making it difficult to identify the desired scene. Furthermore, there are problems with systems that cannot handle changes in the video due to shooting at different distances or angles.
[0712] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0713] In this invention, the server includes: means for receiving video of a scene to be learned and extracting features from the video; means for receiving game video and extracting scenes from the game video that match the features; means for presenting the extracted scenes to a user and allowing the user to determine whether the scene is a specific scene; means for recording the user's determination and updating the learning model to improve extraction accuracy in future runs; means for analyzing the video frame by frame and extracting player movements, positioning, and ball trajectory; means for storing the uploaded video in a storage system; and means for selecting whether the determination result is "correct" or "incorrect." This not only enables automatic extraction of specific scenes with high accuracy, but also enables adaptation to different shooting conditions. Users can quickly and effortlessly confirm the intended scene, achieving efficient analysis work.
[0714] "Means for receiving video of a scene to be learned and extracting features from the video" refers to a device or method in which a server takes in video data containing a specific scene provided by a user, and analyzes and extracts identifiable elements from the video, such as player movements, positioning, and ball trajectory.
[0715] The "means for receiving game footage and extracting scenes from the game footage that match the features" is a mechanism by which the server receives game footage provided by the user, compares it with features learned in advance, and automatically extracts matching parts.
[0716] The "means of presenting the extracted scene to the user and allowing the user to determine whether the scene is a specified scene" refers to a method in which the server displays the video clip extracted as the analysis result to the user through a user interface, allowing the user to check and determine whether the clip is the scene they want.
[0717] "Means for recording the user's judgment results and updating the learning model to improve extraction accuracy in the future" refers to the process in which the server stores the judgment results ("correct" or "incorrect") provided by the user in a database, and retrains and updates the machine learning model based on them to improve the accuracy of subsequent analysis.
[0718] "Means of analyzing video frame by frame to extract player movements, positioning, and ball trajectory" refers to a technology that divides video data into fixed time intervals and analyzes each frame in detail to extract important information such as the movements and positioning of players and the ball.
[0719] The "means for storing uploaded video in a storage system" refers to online storage or a database for temporarily or long-term storage of video data sent by a user to a server.
[0720] The "means for selecting whether the judgment result is 'correct' or 'incorrect'" is a user interface that allows the user to select 'correct' or 'incorrect' as to whether the content of a presented video clip matches a specified scene.
[0721] System Overview
[0722] This invention is a system for efficiently analyzing amateur sports game footage. This system has the function of learning specific scenes and automatically extracting specific scenes from game footage based on the learning data. This allows users to quickly check the intended scenes without any effort, greatly improving the efficiency of the analysis process.
[0723] Hardware and software used
[0724] The hardware used to implement this system includes general servers, terminals, and devices that provide the user interface. The server requires a high-performance CPU and a large amount of storage, and it is recommended to use a cloud service (e.g., AWS or Google Cloud). The terminal can be a smartphone, tablet, or PC.
[0725] The software uses OpenCV for video analysis, TensorFlow or PyTorch for machine learning, and MySQL for database management. The dedicated application allows users to upload recorded video to the server and provides an interface for checking the analysis results.
[0726] Program processing
[0727] Learning Phase
[0728] The user saves the video of the scene they want to extract (for example, a set play during practice) on their device. This video is then uploaded to the server via a dedicated application. The server analyzes the received video frame by frame, extracting features such as player movements, positioning, and ball trajectory, and stores these features as tags in a database.
[0729] Match video analysis phase
[0730] The user saves the game footage on their device and uploads it to the server via a dedicated application. The server receives the game footage and begins video analysis. The game footage is analyzed frame by frame and compared with pre-trained feature vectors. During this analysis process, scenes that match the feature vectors are picked out, organized into a list, and temporarily stored on the server.
[0731] Verification Phase
[0732] The user checks the scene clips extracted from the server through a dedicated application. Each clip is played and judged to be the intended scene. The judgement is made in the dedicated application as "correct" or "incorrect." The user's judgement is sent to the server and recorded in a database. This data is used to update the system's learning model and improve the accuracy of analysis from the next time onwards.
[0733] Specific examples
[0734] The user films a "set play" practice video (e.g., free kick practice) and saves the video on their device. The user then uploads this video to a server using a dedicated application. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball's trajectory and player positioning during a free kick are recorded as tags. The user films a game during the match and uploads the video from their device to the server. The server receives the game video and selects scenes (e.g., free kick scenes) that match the learned features. The server then lists the selected scenes and presents them to the user. The user then uses a dedicated application to review each scene clip from the list and determine whether it is the intended free kick scene. The user then judges whether it is "correct" or "incorrect," and this information is sent to the server. The server uses this data to update the learning model, improving the accuracy of future analysis.
[0735] Prompt Sentence Examples
[0736] "Please explain how to film amateur soccer free kick practice scenes, save them on a device, and upload the footage to a server using a dedicated application. Also, please explain the steps for how the server analyzes and learns from the uploaded footage."
[0737] In this way, the analysis of game footage is automated, significantly reducing the user's workload. Furthermore, the system performs incremental learning, which continuously improves analysis accuracy and enables more accurate scene extraction.
[0738] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0739] Step 1:
[0740] The video of the scene that the user wants to extract (e.g., a set play video during practice) is recorded and saved on the device. The input is a video file containing a specific scene created by the user, and the output is a video file saved on the device. The user uses a smartphone or camera to record the set practice scene and saves the data on the device.
[0741] Step 2:
[0742] The user launches the dedicated application and selects the recorded video file within the application. Then, they click the "Upload" button in the application to upload the video file to the server. The input is the video file saved on the device, and the output is the video data uploaded to the server.
[0743] Step 3:
[0744] The server receives the uploaded video file. It stores the received video file in a storage system (e.g., AWS S3). The input is the uploaded video file, and the output is the stored video data.
[0745] Step 4:
[0746] The server analyzes the video frame by frame. Video processing software such as OpenCV is used for video analysis. Through the analysis, features such as player movements, positioning, and ball trajectory are extracted and stored as tags in a database. The input is the stored video data, and the output is tag data of the extracted features.
[0747] Step 5:
[0748] The user saves the game video on the device. The entire game video is recorded and the data is saved on the device. The input is the game video, and the output is the game video saved on the device. The user also records the game video and saves it on the device.
[0749] Step 6:
[0750] The user uploads game footage saved on their device to the server using a dedicated application. The input is the game footage file saved on the device, and the output is the game footage data uploaded to the server.
[0751] Step 7:
[0752] The server receives and stores the game footage. It analyzes the received footage frame by frame and compares it with pre-trained features. Computer vision technology and machine learning models (e.g., TensorFlow) are used for the matching process. The input is the stored game footage data, and the output is the extraction of scenes that match the features.
[0753] Step 8:
[0754] The server compiles a list of matched scenes and temporarily stores it. The input is the frame data of the matched scenes, and the output is a scene list. These scenes are organized for presentation to the user.
[0755] Step 9:
[0756] The user launches a dedicated application and checks the scene list provided by the server. The user selects the "Check Scene" option within the application and plays a clip of the presented scene. The input is the list provided by the server, and the output is a clip of the scene selected by the user.
[0757] Step 10:
[0758] The user plays each clip and judges whether it is the intended scene as "correct" or "incorrect." This judgment is made on the interface of a dedicated application. The input is the played clip video, and the output is the judgment result, "correct" or "incorrect."
[0759] Step 11:
[0760] The user's judgment results are sent to the server via a dedicated application. The server stores the judgment results in a database and uses this data to retrain and update the machine learning model. The input is the user's judgment results, and the output is the updated learning model. This improves the accuracy of analysis from the next time onwards.
[0761] (Application example 1)
[0762] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0763] This invention relates to a system that efficiently analyzes the manufacturing processes of automated robots and workers used in factories, automatically extracts and evaluates specific work scenes, and improves manufacturing efficiency and detects anomalies. In current manufacturing processes, it is difficult to quickly detect and analyze inefficient operations or abnormalities, which can result in adverse effects on product quality and productivity. To solve this problem, technology is needed that can automatically monitor and analyze manufacturing processes and extract specific work scenes.
[0764] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0765] In this invention, the server includes means for receiving video of a work scene to be learned and extracting features from the video, means for receiving manufacturing process video and extracting scenes from the manufacturing process video that match the features, means for presenting the extracted work scene to a user and allowing the user to determine whether the scene is a predetermined work scene, means for recording the user's determination result and updating the learning model to improve extraction accuracy from the next time onwards, means for analyzing scenes suspected of abnormalities or reduced efficiency and organizing them into a list, and means for generating improvement suggestions based on the analysis results, thereby enabling efficiency improvement in the manufacturing process and early detection of abnormalities.
[0766] "Video of a work scene to be learned" is video data that records the process of a specific work being done in a factory.
[0767] "Features" are data extracted from video data, such as the worker's movements, frequency of tool use, and work time.
[0768] "Manufacturing process video" is video data that records the actual manufacturing process.
[0769] A "matching scene" is a part of the manufacturing process video that has features similar to the learned features.
[0770] "User" refers to a person in a factory who uses the system to analyze the manufacturing process.
[0771] A "predetermined work scene" is a video scene showing the process of a specific work that has been determined in advance.
[0772] The "determination result" is the result of the user's determination as to whether or not the scene is a predetermined work scene.
[0773] A "learning model" is a data processing algorithm that learns features based on video data and improves analysis accuracy.
[0774] "Scenes where abnormalities or reduced efficiency are suspected" are scenes where abnormal operations or reduced work efficiency are observed compared to the normal manufacturing process.
[0775] A "list" is data that organizes and lists situations where abnormalities or reduced efficiency are suspected.
[0776] "Improvement proposals" are specific action plans proposed based on the analysis results to improve the efficiency of manufacturing processes and correct abnormalities.
[0777] The system for implementing this invention starts by recording the manufacturing process of automated robots and workers used in a factory with a camera and uploading the video data to a server. The server then uses dedicated software for analyzing the video data (e.g., OpenCV or TensorFlow) to extract features such as the worker's movements, frequency of tool use, and work time for each frame and stores them in a database.
[0778] Next, actual manufacturing process footage is similarly uploaded to the server and compared with the pre-trained feature values. At this time, the server picks out scenes that are suspected of being abnormal or inefficient and organizes them into a list. The user checks this list using a dedicated device (tablet or smartphone) and determines whether each scene is a specified work scene. The results of this determination are sent to the server, and the learning model is updated, improving extraction accuracy from the next time onwards.
[0779] Furthermore, the server generates improvement proposals based on the analysis results and provides them to users. This enables the efficiency of the manufacturing process and the early detection of abnormalities. In addition, the system has the advantage of being able to continuously improve its analysis accuracy because it learns sequentially.
[0780] Hardware used
[0781] Camera (installed inside the factory)
[0782] Factory Robots
[0783] Dedicated tablet / smartphone
[0784] Servers (e.g., high-performance computing resources on Amazon Web Services (AWS) or Google Cloud Platform (GCP))
[0785] Software used
[0786] Analysis software (e.g., OpenCV (image processing library), TensorFlow (machine learning library))
[0787] Database management system (e.g. MySQL)
[0788] Specific examples
[0789] For example, to analyze a video of a scene in an assembly room in a factory where a robot arm is performing a specific movement and detect a specific movement pattern, the following prompt sentence would be used.
[0790] Prompt Sentence Examples
[0791] "During factory assembly work, identify instances where the robot arm's movement slows down or stops, and compare it with manual work."
[0792] Based on this prompt, the generative AI model analyzes the video data from within the factory, distinguishes between efficient and abnormal operations, and outputs improvement suggestions, allowing users to quickly identify problems in the manufacturing process and take measures.
[0793] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0794] Step 1:
[0795] Users use a dedicated device (tablet or smartphone) to record video of specific tasks within the factory and save it on the device.
[0796] Input: Video data of specific work being done in a factory
[0797] Output: Work video data saved on the device
[0798] Step 2:
[0799] The user uploads the work video data to the server through the terminal.
[0800] Input: Work video data saved on the device
[0801] Output: Work video data uploaded to the server
[0802] Step 3:
[0803] The server receives the uploaded work video data and extracts features using analysis software (e.g., OpenCV or TensorFlow).
[0804] Input: Work video data uploaded to the server
[0805] Output: Extracted features (worker movements, tool usage frequency, work time, etc.)
[0806] Step 4:
[0807] The server stores the extracted features as tags in a database.
[0808] Input: Extracted features
[0809] Output: Tag information stored in the database
[0810] Step 5:
[0811] The user uploads the actual manufacturing process video to the server via the terminal.
[0812] Input: Actual manufacturing process video data
[0813] Output: Manufacturing process video data uploaded to the server
[0814] Step 6:
[0815] The server analyzes the received manufacturing process video data and compares it with the aforementioned features.
[0816] Input: Manufacturing process video data uploaded to the server and tag information stored in the database
[0817] Output: Matching result (scenes that match the features)
[0818] Step 7:
[0819] Identify situations where the server is suspected to be abnormal or inefficient and organize them into a list.
[0820] Input: Matching result
[0821] Output: A list of suspected anomalies and inefficiencies
[0822] Step 8:
[0823] The user checks the list using a dedicated terminal and determines whether each scene is a predetermined work scene.
[0824] Input: List of suspected anomalies or inefficiencies
[0825] Output: Judgment result (whether it is a given work scene or not)
[0826] Step 9:
[0827] The judgment result is sent to the server, and the server updates the learning model.
[0828] Input: Judgment result
[0829] Output: Updated training model
[0830] Step 10:
[0831] The server generates improvement suggestions based on the analysis results and provides them to the user.
[0832] Input: Updated training model and analysis results
[0833] Output: Improvement suggestions
[0834] In this way, specific operations and data processing are carried out at each step, resulting in more efficient manufacturing processes and early detection of abnormalities.
[0835] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0836] System Overview
[0837] This invention is a system for learning and extracting specific scenes for the purpose of efficiently analyzing amateur sports game footage. This system has the ability to learn specific patterns and automatically extract specific scenes from game footage. Furthermore, by combining it with an emotion engine that recognizes user emotions and adjusts the analysis results based on those emotions, it is possible to improve user satisfaction.
[0838] Program processing
[0839] Learning Phase
[0840] 1. The user saves the video of the scene they want to extract (e.g., a set play video during practice) on their device and uploads it to the server via a dedicated application.
[0841] 2. The server receives the uploaded video and begins analysis.
[0842] 3. The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory. These features are stored as tags in a database.
[0843] Match video analysis phase
[0844] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[0845] 2. The server receives the match video and begins video analysis.
[0846] 3. The server analyzes the game footage frame by frame and compares it with pre-trained features. During this analysis process, scenes that match the features are picked out.
[0847] 4. The selected scenes are organized into a list and temporarily saved on the server.
[0848] Emotion Recognition Phase
[0849] 1. When a user uses a dedicated application to check the extracted scene clip, the device analyzes the user's facial expressions, voice, and gestures using an emotion engine.
[0850] 2. The emotion engine recognizes the user's emotions and sends that information to the server.
[0851] 3. The server receives the emotion engine data and dynamically adjusts the content and display method of the presented scene based on the user's emotions.
[0852] 4. The server records changes in the user's emotions and uses this information to improve the way scenes are presented and the accuracy of the analysis content in future sessions.
[0853] Verification Phase
[0854] 1. The user plays each clip and judges whether it is the intended scene. This judgment is made by a dedicated application, with a "correct" or "incorrect" result.
[0855] 2. The user's judgment results are sent to the server and recorded in a database. This data is used to update the system's learning model and improve the accuracy of analysis in future.
[0856] Specific examples
[0857] 1. The user takes a video of a "set play" practice (e.g., free kick practice) and saves the video on their device.
[0858] 2. The user uploads this video to the server using a dedicated application.
[0859] 3. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball trajectory and player positioning during a free kick are recorded as tags.
[0860] 4. The user takes video of the game and uploads it from their device to the server.
[0861] 5. The server receives the game footage and picks out scenes that match the learned features (e.g., free kick scenes).
[0862] 6. The server lists the scenes picked up and presents them to the user.
[0863] 7. When the user uses a dedicated application to check clips for each scene from the list, the device analyzes the user's facial expressions, voice, and gestures, and the emotion engine recognizes the user's emotions.
[0864] 8. The server dynamically adjusts how the clip is presented based on the user's emotions, for example, adjusting the playback speed of the clip if a positive emotion is detected, or revisiting a section of the clip if a negative emotion is detected.
[0865] 9. The user marks each clip as "true" or "false," and the result is sent to the server.
[0866] 10. The server records the judgment data and emotion data from the user, updates the system's learning model, and improves the accuracy of analysis from the next time onwards.
[0867] In this way, game video analysis is automated and dynamic scene adjustments based on the user's emotions are made possible, providing a more intuitive and satisfying video analysis experience. The system continues to learn and improves its analysis accuracy, resulting in an increasingly rich analysis experience.
[0868] The processing flow will be explained below.
[0869] Learning Phase
[0870] Step 1:
[0871] The video of the scene the user wants to extract (e.g., a set play video during practice) is saved on the device.
[0872] Step 2:
[0873] The user launches the dedicated application and selects the video of the scene they want to extract.
[0874] Step 3:
[0875] The terminal uploads the selected video to the server.
[0876] Step 4:
[0877] The server receives the uploaded video and begins analysis.
[0878] Step 5:
[0879] The server analyzes the video frame by frame and extracts features such as players' movements, positioning, and ball trajectory.
[0880] Step 6:
[0881] The server stores the extracted features as tags in a database.
[0882] Match video analysis phase
[0883] Step 1:
[0884] The user saves the game video on the device.
[0885] Step 2:
[0886] The user launches the dedicated application and selects the game footage.
[0887] Step 3:
[0888] The terminal uploads the selected game video to the server.
[0889] Step 4:
[0890] The server prepares to analyze the received game footage.
[0891] Step 5:
[0892] The server analyzes the game footage frame by frame and compares it with pre-learned features.
[0893] Step 6:
[0894] The server picks out scenes that match the matched features.
[0895] Step 7:
[0896] The server organizes the picked scenes into a list and temporarily saves them.
[0897] Emotion Recognition Phase
[0898] Step 1:
[0899] The user launches the dedicated application and checks the scene clips extracted from the server.
[0900] Step 2:
[0901] The device uses an emotion engine to analyze the user's emotional information, such as facial expressions, voice, and gestures.
[0902] Step 3:
[0903] The emotion engine recognizes the user's emotions and sends the information to the server.
[0904] Step 4:
[0905] The server receives the emotion engine data and dynamically adjusts the presentation method and content of the extracted scene based on the user's emotion.
[0906] Step 5:
[0907] The server records changes in the user's emotions and uses this information to improve the accuracy of scene presentations and analysis content in future sessions.
[0908] Verification Phase
[0909] Step 1:
[0910] The user plays each clip and determines whether it is the intended scene.
[0911] Step 2:
[0912] The user selects "true" or "false" for each clip in a dedicated application.
[0913] Step 3:
[0914] The user's decision result is sent to the server.
[0915] Step 4:
[0916] The server stores the judgment results and emotion recognition data in a database.
[0917] Step 5:
[0918] The server uses the stored data to update the learning model and improve the accuracy of analysis from the next time onwards.
[0919] Specific examples
[0920] Learning Phase
[0921] Step 1:
[0922] The user saves a "set play" practice video (e.g., free kick practice) on the device.
[0923] Step 2:
[0924] The user launches the dedicated application, selects a practice video, and uploads it to the server.
[0925] Step 3:
[0926] The device sends the video to the server.
[0927] Step 4:
[0928] The server receives the video and prepares it for analysis.
[0929] Step 5:
[0930] The server analyzes the video frame by frame and extracts features such as player movements, ball trajectory, and player positioning.
[0931] Step 6:
[0932] The server registers the extracted features in a database.
[0933] Match video analysis phase
[0934] Step 1:
[0935] The user saves the game video on the device.
[0936] Step 2:
[0937] The user launches the dedicated application, selects the game footage, and uploads it to the server.
[0938] Step 3:
[0939] The device sends the video to the server.
[0940] Step 4:
[0941] The server receives the match footage and prepares it for analysis.
[0942] Step 5:
[0943] The server analyzes the game footage frame by frame and compares it with the learned features.
[0944] Step 6:
[0945] The server picks out scenes that match the matched features (e.g., free kick scenes).
[0946] Step 7:
[0947] The server creates a list of the scenes it picks up and temporarily saves them.
[0948] Emotion Recognition Phase
[0949] Step 1:
[0950] The user starts the dedicated application and checks the clips of the scenes listed from the server.
[0951] Step 2:
[0952] The device analyzes the user's facial expressions, voice, gestures, etc. using an emotion engine.
[0953] Step 3:
[0954] The emotion engine recognizes the user's emotion and transmits the emotion information to the server.
[0955] Step 4:
[0956] The server receives the emotion engine data and dynamically adjusts how the clip is displayed based on the user's emotion.
[0957] Step 5:
[0958] The server records the user's emotional fluctuations and uses this information to improve the way scenes are presented and the accuracy of the analysis content from next time onwards.
[0959] Verification Phase
[0960] Step 1:
[0961] The user plays each clip and checks whether it is the intended scene.
[0962] Step 2:
[0963] The user selects "true" or "false" for each clip and sends the result to the server.
[0964] Step 3:
[0965] The server stores the judgment results and emotion recognition data in a database.
[0966] Step 4:
[0967] The server uses the stored data to update the learning model and improve analysis accuracy.
[0968] This series of processes automates the analysis of match footage and enables dynamic scene adjustments based on the user's emotions, providing a more intuitive and satisfying video analysis experience. The system continues to learn and improves its analysis accuracy, resulting in an increasingly rich analysis experience.
[0969] Example 2
[0970] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0971] When analyzing amateur sports game footage, it is difficult to efficiently and accurately extract specific scenes. Furthermore, to improve user satisfaction with the information obtained from video analysis, it is necessary to dynamically adjust the video based on the user's emotions. However, current systems lack the functionality to recognize the user's emotions and dynamically adjust the video presentation based on those emotions. As a result, users may not obtain the information they expect, potentially reducing the effectiveness of video analysis. A system that can solve these issues is needed.
[0972] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0973] In this invention, the server includes means for receiving video of scenes to be learned and extracting features from the video, means for receiving game video and extracting scenes from the game video that match the features, means for presenting the extracted scenes to a user and allowing the user to determine whether the scenes are predetermined scenes, means for recording the user's determination result and updating the learning model to improve extraction accuracy in the future, and means for recognizing the user's emotions and dynamically adjusting the content and presentation method of the extracted scenes based on the recognized emotions. This enables efficient and accurate scene extraction from game video and dynamic adjustment of the video in accordance with the user's emotions, thereby achieving video analysis with high user satisfaction.
[0974] Below are definitions of important words.
[0975] "Video of a scene to be learned" refers to video data that includes a specific scene that the user wants to extract.
[0976] "Features" refers to information extracted from video data that is necessary to identify specific scenes, such as player movements, positioning, and ball trajectory.
[0977] "Game footage" refers to video data that records the progress of an actual sports game.
[0978] An "emotion engine" refers to software or hardware that analyzes a user's facial expressions, voice, and gestures to recognize emotions.
[0979] A "learning model" refers to the algorithms and data structures that allow a system to continuously learn based on past data.
[0980] "Server" refers to a computer system that processes and stores video data and analysis data uploaded by users and performs various analyses.
[0981] "Terminal" refers to the device (e.g., smartphone, tablet, PC, etc.) that a user uses to store video and operate a dedicated application.
[0982] "Tag" refers to a label or metadata that is added to identify a specific feature in a video.
[0983] "Extraction" refers to extracting features or scenes that match predetermined conditions from video data.
[0984] "Dynamic adjustment" refers to changing the way the video is displayed in real time based on the user's emotions and other variables.
[0985] This allows the technical scope of the invention to be more clearly defined.
[0986] This invention is a system for efficiently analyzing amateur sports game footage, and has the ability to learn and extract specific scenes. This system learns specific patterns and automatically extracts specific scenes from game footage. Furthermore, by combining it with an emotion engine that recognizes user emotions and adjusts the analysis results based on those emotions, it is possible to improve user satisfaction.
[0987] The main hardware and software components for implementing this system are as follows:
[0988] 1. Device: A device used by a user to store video and operate dedicated applications. Examples include smartphones, tablets, and PCs.
[0989] 2. Server: A computer system that processes and stores video data and analysis data uploaded by users and performs various analyses. It is responsible for video analysis, database management, and execution of learning models.
[0990] 3. Emotion engine: Software or hardware that analyzes the user's facial expressions, voice, and gestures to recognize emotions.
[0991] The specific steps for implementing the system are as follows:
[0992] Learning Phase
[0993] 1. The user takes a video of the scene they want to extract and saves it on their device.
[0994] 2. The user uploads the video to the server through a dedicated application.
[0995] 3. The server receives the video and analyzes it frame by frame to extract features such as player movements, positioning, and ball trajectory.
[0996] Match video analysis phase
[0997] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[0998] 2. The server receives the game footage, analyzes each frame, and compares it with pre-trained features.
[0999] 3. The server picks up matching scenes, organizes them in a list format, and temporarily saves them.
[1000] Emotion Recognition Phase
[1001] 1. The user uses a dedicated application to view clips of the scenes selected here.
[1002] 2. The device sends the user's facial expressions, voice, and gestures to the emotion engine.
[1003] 3. The emotion engine generates emotion data and sends it to the server.
[1004] 4. The server dynamically adjusts the content and presentation of the clip based on the emotional data.
[1005] Verification Phase
[1006] 1. The user judges each clip as "true" or "false."
[1007] 2. The user's judgment result is sent to the server and recorded in the database.
[1008] 3. The server uses the judgment results to update the learning model and improve the accuracy of analysis from the next time onwards.
[1009] For example, if a user films and uploads a free-kick practice video, the server analyzes the ball's trajectory and player positions to select free-kick scenes from game footage. Then, when the user plays back the video, the system can dynamically adjust the playback speed and content of the clip based on the user's emotions.
[1010] Prompt Sentence Examples
[1011] 1. "Please explain the scenario where a user uploads a video of themselves practicing free kicks."
[1012] 2. "Please explain how to analyze frames from game footage."
[1013] 3. "Please explain how the emotion engine recognizes the user's emotions and how the server uses them."
[1014] This system enables efficient and accurate extraction of specific scenes, and also realizes dynamic adjustment of the video according to the user's emotions, thereby providing video analysis with high user satisfaction.
[1015] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1016] Learning Phase
[1017] Step 1:
[1018] The user takes a video of the scene they want to extract and saves it on their device.
[1019] Input: Video files captured by the camera.
[1020] Output: Video data saved on the device.
[1021] Step 2:
[1022] The user opens the dedicated application, selects the saved video file, and uploads it to the server.
[1023] Input: User specified video file.
[1024] Output: Video data received by the server.
[1025] Step 3:
[1026] The server starts a program to analyze the received video file.
[1027] Input: Uploaded video file.
[1028] Output: Video data divided into frames.
[1029] Step 4:
[1030] The server extracts features such as player movements, positioning, and ball trajectory for each frame.
[1031] Input: Video data divided into frames.
[1032] Output: Extracted feature data (player movements and positions, ball trajectory, etc.).
[1033] Step 5:
[1034] The server stores the extracted features as tags in a database.
[1035] Input: Extracted feature data.
[1036] Output: Tag information stored in the database.
[1037] Match video analysis phase
[1038] Step 1:
[1039] Users save game footage on their devices and upload it to the server via a dedicated application.
[1040] Input: User-specified game video file.
[1041] Output: Match video data received by the server.
[1042] Step 2:
[1043] The server receives the game footage and starts a program that analyzes it frame by frame.
[1044] Input: Uploaded match video file.
[1045] Output: Match video data divided into frames.
[1046] Step 3:
[1047] The server compares the features with those learned in advance and picks out scenes with matching features.
[1048] Input: Game video data divided into frames and pre-trained feature data.
[1049] Output: Frame number and time information of the matching scenes.
[1050] Step 4:
[1051] The server organizes the picked-up scenes in list format and temporarily saves them.
[1052] Input: Frame number and time information of the matching scene.
[1053] Output: List of scene information and its temporary storage.
[1054] Emotion Recognition Phase
[1055] Step 1:
[1056] The user can use a dedicated application to check clips of the scenes picked up here.
[1057] Input: Listed scene information.
[1058] Output: The clip of the scene being played.
[1059] Step 2:
[1060] The device transmits the user's facial expressions, voice, and gestures to the emotion engine.
[1061] Input: User facial, voice, and gesture data.
[1062] Output: Data sent to the emotion engine.
[1063] Step 3:
[1064] The emotion engine generates emotion data and sends it to the server.
[1065] Input: User facial, voice, and gesture data.
[1066] Output: The generated emotion data.
[1067] Step 4:
[1068] The server dynamically adjusts the content and presentation method of the clip based on the emotional data received.
[1069] Input: Emotion data and playing clip information.
[1070] Output: The playback speed and display of the adjusted clip.
[1071] Verification Phase
[1072] Step 1:
[1073] The user marks each clip as "true" or "false."
[1074] Input: The clip that was played.
[1075] Output: User's decision result.
[1076] Step 2:
[1077] The user's judgment result is sent to the server and recorded in a database.
[1078] Input: User's decision result.
[1079] Output: Verdict data recorded in a database.
[1080] Step 3:
[1081] The server uses the judgment results to update the learning model, improving the accuracy of analysis from the next time onwards.
[1082] Input: Recorded adjudication data.
[1083] Output: Updated training model.
[1084] (Application example 2)
[1085] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1086] When analyzing video of sports games, advertisements, etc., there is a need to efficiently extract specific scenes and dynamically adjust the display content based on the user's emotions. However, conventional technologies have had difficulty recognizing the user's emotions in real time and optimally displaying scenes and delivering advertisements in accordance with those emotions. The present invention aims to solve this problem and provide more intuitive and satisfying video analysis and advertisement delivery.
[1087] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1088] In this invention, the server includes means for receiving video of scenes to be learned and extracting features from the video, means for receiving game video and extracting scenes from the game video that match the features, and means for recognizing the user's emotions while watching and dynamically adjusting the content and display method of the extracted scenes based on the emotions, thereby enabling scene display and advertisement delivery in accordance with the user's emotions.
[1089] "Video of the scene to be learned" refers to video that the system uses to extract features in advance.
[1090] "Features" are data such as player movements, positioning, and ball trajectory that the system extracts from the video.
[1091] "Game footage" refers to footage of an actual sport or event being played.
[1092] "User" means any person or entity that uses the System.
[1093] "Means for adjusting based on emotions" is a mechanism that analyzes the user's emotions and dynamically changes the content and method of displaying the video based on the results.
[1094] "Dynamic adjustment" refers to changing the content and display method of the video in real time in response to changes in the user's emotions.
[1095] A "server" refers to a computer or network system that performs the main processing of a system.
[1096] System Overview
[1097] This invention is a system for efficiently analyzing sports game footage and advertising videos. It is characterized by recognizing the user's emotions while watching and dynamically adjusting the content and display method of the video based on those emotions. This system extracts features from the video of the scene being studied and analyzes the game footage and advertising videos in real time to provide the user with an optimal experience.
[1098] Program processing
[1099] Learning Phase
[1100] The server receives the video of the scene the user wants to extract (the learning video) and extracts features from the video. Specifically, it analyzes the video frame by frame and uses a generative AI model to extract data such as player movements, positioning, and ball trajectory. These features are then stored as tags in a database.
[1101] Match video analysis phase
[1102] Users save game footage or advertising footage on their devices and upload it to the server via a dedicated application. The server analyzes the received footage and compares it with pre-trained features to automatically select scenes that match. These scenes are organized into a list and temporarily saved.
[1103] Emotion Recognition Phase
[1104] When a user reviews an extracted scene using a dedicated application, the device analyzes the user's facial expressions, voice, and gestures using an emotion engine. The server receives the data sent from the emotion engine and dynamically adjusts how the scene is displayed based on the user's emotions. If a positive emotion is recognized, the playback speed of the clip is adjusted, and if a negative emotion is recognized, the clip is reviewed.
[1105] Hardware and software used
[1106] Hardware: High-performance servers, user devices (smartphones, tablets, PCs, etc.)
[1107] Software: Generative AI models (Keras, TensorFlow, OpenCV), emotion engine, database management system
[1108] Specific examples
[1109] For example, the present invention can be implemented as a system for delivering advertisements based on user emotions in the advertising industry. The server analyzes the advertisement video being viewed by the user in real time, recognizes the user's emotions while viewing, and dynamically selects the next advertisement to be delivered based on the emotions.
[1110] Examples of prompts:
[1111] "Please build a system that analyzes the advertisements that users watch in real time, recognizes their emotions from their facial expressions while watching, and dynamically selects the next advertisement to be delivered."
[1112] This makes it possible to deliver optimal video and advertisements according to the user's emotions, improving the viewing experience and maximizing the effectiveness of advertising.
[1113] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1114] Step 1:
[1115] Input: The user saves the video of the scene to be learned on the device.
[1116] How it works: The user uses a dedicated application to upload the video to be studied to the server.
[1117] Output: The learning video is saved on the server.
[1118] Step 2:
[1119] Input: The training video received by the server.
[1120] How it works: The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory, using generative AI models (Keras, TensorFlow).
[1121] Output: The extracted features are stored as tags in a database.
[1122] Step 3:
[1123] Input: The user saves game footage or advertising footage to the device.
[1124] How it works: Users use a dedicated application to upload game footage and advertising footage to the server.
[1125] Output: Match footage and advertising footage are saved on the server.
[1126] Step 4:
[1127] Input: Game footage and advertising footage received by the server, as well as pre-trained features.
[1128] How it works: The server analyzes the received video frame by frame and compares it with pre-trained features to pick out matching scenes. It then uses a generative AI model to confirm feature matches.
[1129] Output: Matching scenes are organized into a list and temporarily stored on the server.
[1130] Step 5:
[1131] Input: The user checks the extracted scene clips in a dedicated application.
[1132] How it works: While the user plays the clip, the device uses the camera and microphone to input the user's facial expressions, voice, and gestures into the emotion engine in real time.
[1133] Output: The emotion engine extracts the user's emotion data and sends it to the server.
[1134] Step 6:
[1135] Input: Emotion data received by the server, and a list of extracted scenes.
[1136] How it works: The server dynamically adjusts how the clip is presented based on the user's emotional data, for example adjusting the clip's playback speed if a positive emotion is detected, or revisiting a section of the clip if a negative emotion is detected.
[1137] Output: The adjusted scene is displayed to the user in real time.
[1138] Step 7:
[1139] Input: The user judges each clip as "true" or "false."
[1140] How it works: The user uses a dedicated application to send the results of their assessment of the clip to the server.
[1141] Output: The server records the user's judgment results in a database and updates the learning model to improve analysis accuracy in future.
[1142] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1143] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1144] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1145] [Third embodiment]
[1146] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1147] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1148] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1149] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1150] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1151] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1152] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1153] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1154] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1155] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1156] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1157] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1158] System Overview
[1159] This invention is a system for efficiently analyzing amateur sports game footage. This system has the function of learning specific scenes and automatically extracting specific scenes from game footage based on the learning data. This allows users to quickly check the intended scenes without any effort, greatly improving the efficiency of the analysis process.
[1160] Program processing
[1161] Learning Phase
[1162] 1. The user saves the video of the scene they want to extract (e.g., a set play video during practice) on their device and uploads it to the server via a dedicated application.
[1163] 2. The server receives the uploaded video and begins analysis.
[1164] 3. The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory. These features are stored as tags in a database.
[1165] Match video analysis phase
[1166] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[1167] 2. The server receives the match video and begins video analysis.
[1168] 3. The server analyzes the game footage frame by frame and compares it with pre-trained features. During this analysis process, scenes that match the features are picked out.
[1169] 4. The selected scenes are organized into a list and temporarily saved on the server.
[1170] Verification Phase
[1171] 1. The user checks the scene clips extracted from the server through a dedicated application.
[1172] 2. The user plays each clip and judges whether it is the intended scene. This judgment is made by a dedicated application, with a "correct" or "incorrect" result.
[1173] 3. The user's judgment results are sent to the server and recorded in a database. This data is used to update the system's learning model and improve analysis accuracy in future analyses.
[1174] Specific examples
[1175] 1. The user takes a video of a "set play" practice (e.g., free kick practice) and saves the video on their device.
[1176] 2. The user uploads this video to the server using a dedicated application.
[1177] 3. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball trajectory and player positioning during a free kick are recorded as tags.
[1178] 4. The user takes video of the game and uploads it from their device to the server.
[1179] 5. The server receives the game footage and picks out scenes that match the learned features (e.g., free kick scenes).
[1180] 6. The server lists the scenes picked up and presents them to the user.
[1181] 7. The user uses a dedicated application to check the clips of each scene from the list and determine whether they are the intended free kick scenes.
[1182] 8. The user judges whether the result is "correct" or "incorrect," and this information is sent to the server. The server uses this data to update the learning model and improve the accuracy of future analyses.
[1183] In this way, the analysis of game footage is automated, significantly reducing the user's workload. Furthermore, the system performs incremental learning, which continuously improves analysis accuracy and enables more accurate scene extraction.
[1184] The processing flow will be explained below.
[1185] Learning Phase
[1186] Step 1:
[1187] The video of the scene the user wants to extract (e.g., a set play video during practice) is saved on the device.
[1188] Step 2:
[1189] The user launches the dedicated application and selects the video of the scene they want to extract.
[1190] Step 3:
[1191] The terminal uploads the selected video to the server.
[1192] Step 4:
[1193] The server prepares to analyze the received video.
[1194] Step 5:
[1195] The server analyzes the video frame by frame and extracts features such as players' movements, positioning, and ball trajectory.
[1196] Step 6:
[1197] The server stores the extracted features as tags in a database.
[1198] Match video analysis phase
[1199] Step 1:
[1200] The user saves the game video on the device.
[1201] Step 2:
[1202] The user launches the dedicated application and selects the game footage.
[1203] Step 3:
[1204] The terminal uploads the selected game video to the server.
[1205] Step 4:
[1206] The server prepares to analyze the received game footage.
[1207] Step 5:
[1208] The server analyzes the game footage frame by frame and compares it with pre-learned features.
[1209] Step 6:
[1210] The server picks out scenes that match the matched features.
[1211] Step 7:
[1212] The server organizes the picked scenes into a list and temporarily saves them.
[1213] Verification Phase
[1214] Step 1:
[1215] The user launches the dedicated application and checks the scene clips extracted from the server.
[1216] Step 2:
[1217] The user plays each clip and determines whether it is the intended scene.
[1218] Step 3:
[1219] The user selects "true" or "false" for each clip in a dedicated application.
[1220] Step 4:
[1221] The user's decision result is sent to the server.
[1222] Step 5:
[1223] The server records the received judgment result in a database.
[1224] Step 6:
[1225] The server uses the recorded data to update the learning model, improving the accuracy of analysis from the next time onwards.
[1226] Specific examples
[1227] Learning Phase
[1228] Step 1:
[1229] The user saves a "set play" practice video (e.g., free kick practice) on the device.
[1230] Step 2:
[1231] The user launches the dedicated application, selects a practice video, and uploads it to the server.
[1232] Step 3:
[1233] The device sends the video to the server.
[1234] Step 4:
[1235] The server receives the video and prepares it for analysis.
[1236] Step 5:
[1237] The server analyzes the video frame by frame and extracts features such as player movements, ball trajectory, and player positioning.
[1238] Step 6:
[1239] The server registers the extracted features in a database.
[1240] Match video analysis phase
[1241] Step 1:
[1242] The user saves the game video on the device.
[1243] Step 2:
[1244] The user launches the dedicated application, selects the game footage, and uploads it to the server.
[1245] Step 3:
[1246] The device sends the video to the server.
[1247] Step 4:
[1248] The server receives the match footage and prepares it for analysis.
[1249] Step 5:
[1250] The server analyzes the game footage frame by frame and compares it with the learned features.
[1251] Step 6:
[1252] The server picks out scenes that match the matched features (e.g., free kick scenes).
[1253] Step 7:
[1254] The server creates a list of the scenes it picks up and temporarily saves them.
[1255] Verification Phase
[1256] Step 1:
[1257] The user starts the dedicated application and checks the clip of the scene notified by the server.
[1258] Step 2:
[1259] The user plays each clip and checks whether it is the intended scene.
[1260] Step 3:
[1261] The user uses a dedicated application to make a judgment ("correct" or "incorrect") about each clip.
[1262] Step 4:
[1263] The user's decision result is sent to the server.
[1264] Step 5:
[1265] The server stores the received judgment results in a database.
[1266] Step 6:
[1267] The server uses the stored data to update the analysis model and improve accuracy from the next time onwards.
[1268] This automates the analysis of game footage, allowing users to efficiently review the scenes they intended. The system continues to learn, improving its analysis accuracy and providing an increasingly rich analysis experience.
[1269] Example 1
[1270] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1271] In conventional amateur sports game video analysis systems, extracting specific scenes is a manual process that requires a great deal of time and effort. Furthermore, the accuracy of scene analysis is low, making it difficult to identify the desired scene. Furthermore, there are problems with systems that cannot handle changes in the video due to shooting at different distances or angles.
[1272] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1273] In this invention, the server includes: means for receiving video of a scene to be learned and extracting features from the video; means for receiving game video and extracting scenes from the game video that match the features; means for presenting the extracted scenes to a user and allowing the user to determine whether the scene is a specific scene; means for recording the user's determination and updating the learning model to improve extraction accuracy in future runs; means for analyzing the video frame by frame and extracting player movements, positioning, and ball trajectory; means for storing the uploaded video in a storage system; and means for selecting whether the determination result is "correct" or "incorrect." This not only enables automatic extraction of specific scenes with high accuracy, but also enables adaptation to different shooting conditions. Users can quickly and effortlessly confirm the intended scene, achieving efficient analysis work.
[1274] "Means for receiving video of a scene to be learned and extracting features from the video" refers to a device or method in which a server takes in video data containing a specific scene provided by a user, and analyzes and extracts identifiable elements from the video, such as player movements, positioning, and ball trajectory.
[1275] The "means for receiving game footage and extracting scenes from the game footage that match the features" is a mechanism by which the server receives game footage provided by the user, compares it with features learned in advance, and automatically extracts matching parts.
[1276] The "means of presenting the extracted scene to the user and allowing the user to determine whether the scene is a specified scene" refers to a method in which the server displays the video clip extracted as the analysis result to the user through a user interface, allowing the user to check and determine whether the clip is the scene they want.
[1277] "Means for recording the user's judgment results and updating the learning model to improve extraction accuracy in the future" refers to the process in which the server stores the judgment results ("correct" or "incorrect") provided by the user in a database, and retrains and updates the machine learning model based on them to improve the accuracy of subsequent analysis.
[1278] "Means of analyzing video frame by frame to extract player movements, positioning, and ball trajectory" refers to a technology that divides video data into fixed time intervals and analyzes each frame in detail to extract important information such as the movements and positioning of players and the ball.
[1279] The "means for storing uploaded video in a storage system" refers to online storage or a database for temporarily or long-term storage of video data sent by a user to a server.
[1280] The "means for selecting whether the judgment result is 'correct' or 'incorrect'" is a user interface that allows the user to select 'correct' or 'incorrect' as to whether the content of a presented video clip matches a specified scene.
[1281] System Overview
[1282] This invention is a system for efficiently analyzing amateur sports game footage. This system has the function of learning specific scenes and automatically extracting specific scenes from game footage based on the learning data. This allows users to quickly check the intended scenes without any effort, greatly improving the efficiency of the analysis process.
[1283] Hardware and software used
[1284] The hardware used to implement this system includes general servers, terminals, and devices that provide the user interface. The server requires a high-performance CPU and a large amount of storage, and it is recommended to use a cloud service (e.g., AWS or Google Cloud). The terminal can be a smartphone, tablet, or PC.
[1285] The software uses OpenCV for video analysis, TensorFlow or PyTorch for machine learning, and MySQL for database management. The dedicated application allows users to upload recorded video to the server and provides an interface for checking the analysis results.
[1286] Program processing
[1287] Learning Phase
[1288] The user saves the video of the scene they want to extract (for example, a set play during practice) on their device. This video is then uploaded to the server via a dedicated application. The server analyzes the received video frame by frame, extracting features such as player movements, positioning, and ball trajectory, and stores these features as tags in a database.
[1289] Match video analysis phase
[1290] The user saves the game footage on their device and uploads it to the server via a dedicated application. The server receives the game footage and begins video analysis. The game footage is analyzed frame by frame and compared with pre-trained feature vectors. During this analysis process, scenes that match the feature vectors are picked out, organized into a list, and temporarily stored on the server.
[1291] Verification Phase
[1292] The user checks the scene clips extracted from the server through a dedicated application. Each clip is played and judged to be the intended scene. The judgement is made in the dedicated application as "correct" or "incorrect." The user's judgement is sent to the server and recorded in a database. This data is used to update the system's learning model and improve the accuracy of analysis from the next time onwards.
[1293] Specific examples
[1294] The user films a "set play" practice video (e.g., free kick practice) and saves the video on their device. The user then uploads this video to a server using a dedicated application. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball's trajectory and player positioning during a free kick are recorded as tags. The user films a game during the match and uploads the video from their device to the server. The server receives the game video and selects scenes (e.g., free kick scenes) that match the learned features. The server then lists the selected scenes and presents them to the user. The user then uses a dedicated application to review each scene clip from the list and determine whether it is the intended free kick scene. The user then judges whether it is "correct" or "incorrect," and this information is sent to the server. The server uses this data to update the learning model, improving the accuracy of future analysis.
[1295] Prompt Sentence Examples
[1296] "Please explain how to film amateur soccer free kick practice scenes, save them on a device, and upload the footage to a server using a dedicated application. Also, please explain the steps for how the server analyzes and learns from the uploaded footage."
[1297] In this way, the analysis of game footage is automated, significantly reducing the user's workload. Furthermore, the system performs incremental learning, which continuously improves analysis accuracy and enables more accurate scene extraction.
[1298] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1299] Step 1:
[1300] The video of the scene that the user wants to extract (e.g., a set play video during practice) is recorded and saved on the device. The input is a video file containing a specific scene created by the user, and the output is a video file saved on the device. The user uses a smartphone or camera to record the set practice scene and saves the data on the device.
[1301] Step 2:
[1302] The user launches the dedicated application and selects the recorded video file within the application. Then, they click the "Upload" button in the application to upload the video file to the server. The input is the video file saved on the device, and the output is the video data uploaded to the server.
[1303] Step 3:
[1304] The server receives the uploaded video file. It stores the received video file in a storage system (e.g., AWS S3). The input is the uploaded video file, and the output is the stored video data.
[1305] Step 4:
[1306] The server analyzes the video frame by frame. Video processing software such as OpenCV is used for video analysis. Through the analysis, features such as player movements, positioning, and ball trajectory are extracted and stored as tags in a database. The input is the stored video data, and the output is tag data of the extracted features.
[1307] Step 5:
[1308] The user saves the game video on the device. The entire game video is recorded and the data is saved on the device. The input is the game video, and the output is the game video saved on the device. The user also records the game video and saves it on the device.
[1309] Step 6:
[1310] The user uploads game footage saved on their device to the server using a dedicated application. The input is the game footage file saved on the device, and the output is the game footage data uploaded to the server.
[1311] Step 7:
[1312] The server receives and stores the game footage. It analyzes the received footage frame by frame and compares it with pre-trained features. Computer vision technology and machine learning models (e.g., TensorFlow) are used for the matching process. The input is the stored game footage data, and the output is the extraction of scenes that match the features.
[1313] Step 8:
[1314] The server compiles a list of matched scenes and temporarily stores it. The input is the frame data of the matched scenes, and the output is a scene list. These scenes are organized for presentation to the user.
[1315] Step 9:
[1316] The user launches a dedicated application and checks the scene list provided by the server. The user selects the "Check Scene" option within the application and plays a clip of the presented scene. The input is the list provided by the server, and the output is a clip of the scene selected by the user.
[1317] Step 10:
[1318] The user plays each clip and judges whether it is the intended scene as "correct" or "incorrect." This judgment is made on the interface of a dedicated application. The input is the played clip video, and the output is the judgment result, "correct" or "incorrect."
[1319] Step 11:
[1320] The user's judgment results are sent to the server via a dedicated application. The server stores the judgment results in a database and uses this data to retrain and update the machine learning model. The input is the user's judgment results, and the output is the updated learning model. This improves the accuracy of analysis from the next time onwards.
[1321] (Application example 1)
[1322] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1323] This invention relates to a system that efficiently analyzes the manufacturing processes of automated robots and workers used in factories, automatically extracts and evaluates specific work scenes, and improves manufacturing efficiency and detects anomalies. In current manufacturing processes, it is difficult to quickly detect and analyze inefficient operations or abnormalities, which can result in adverse effects on product quality and productivity. To solve this problem, technology is needed that can automatically monitor and analyze manufacturing processes and extract specific work scenes.
[1324] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1325] In this invention, the server includes means for receiving video of a work scene to be learned and extracting features from the video, means for receiving manufacturing process video and extracting scenes from the manufacturing process video that match the features, means for presenting the extracted work scene to a user and allowing the user to determine whether the scene is a predetermined work scene, means for recording the user's determination result and updating the learning model to improve extraction accuracy from the next time onwards, means for analyzing scenes suspected of abnormalities or reduced efficiency and organizing them into a list, and means for generating improvement suggestions based on the analysis results, thereby enabling efficiency improvement in the manufacturing process and early detection of abnormalities.
[1326] "Video of a work scene to be learned" is video data that records the process of a specific work being done in a factory.
[1327] "Features" are data extracted from video data, such as the worker's movements, frequency of tool use, and work time.
[1328] "Manufacturing process video" is video data that records the actual manufacturing process.
[1329] A "matching scene" is a part of the manufacturing process video that has features similar to the learned features.
[1330] "User" refers to a person in a factory who uses the system to analyze the manufacturing process.
[1331] A "predetermined work scene" is a video scene showing the process of a specific work that has been determined in advance.
[1332] The "determination result" is the result of the user's determination as to whether or not the scene is a predetermined work scene.
[1333] A "learning model" is a data processing algorithm that learns features based on video data and improves analysis accuracy.
[1334] "Scenes where abnormalities or reduced efficiency are suspected" are scenes where abnormal operations or reduced work efficiency are observed compared to the normal manufacturing process.
[1335] A "list" is data that organizes and lists situations where abnormalities or reduced efficiency are suspected.
[1336] "Improvement proposals" are specific action plans proposed based on the analysis results to improve the efficiency of manufacturing processes and correct abnormalities.
[1337] The system for implementing this invention starts by recording the manufacturing process of automated robots and workers used in a factory with a camera and uploading the video data to a server. The server then uses dedicated software for analyzing the video data (e.g., OpenCV or TensorFlow) to extract features such as the worker's movements, frequency of tool use, and work time for each frame and stores them in a database.
[1338] Next, actual manufacturing process footage is similarly uploaded to the server and compared with the pre-trained feature values. At this time, the server picks out scenes that are suspected of being abnormal or inefficient and organizes them into a list. The user checks this list using a dedicated device (tablet or smartphone) and determines whether each scene is a specified work scene. The results of this determination are sent to the server, and the learning model is updated, improving extraction accuracy from the next time onwards.
[1339] Furthermore, the server generates improvement proposals based on the analysis results and provides them to users. This enables the efficiency of the manufacturing process and the early detection of abnormalities. In addition, the system has the advantage of being able to continuously improve its analysis accuracy because it learns sequentially.
[1340] Hardware used
[1341] Camera (installed inside the factory)
[1342] Factory Robots
[1343] Dedicated tablet / smartphone
[1344] Servers (e.g., high-performance computing resources on Amazon Web Services (AWS) or Google Cloud Platform (GCP))
[1345] Software used
[1346] Analysis software (e.g., OpenCV (image processing library), TensorFlow (machine learning library))
[1347] Database management system (e.g. MySQL)
[1348] Specific examples
[1349] For example, to analyze a video of a scene in an assembly room in a factory where a robot arm is performing a specific movement and detect a specific movement pattern, the following prompt sentence would be used.
[1350] Prompt Sentence Examples
[1351] "During factory assembly work, identify instances where the robot arm's movement slows down or stops, and compare it with manual work."
[1352] Based on this prompt, the generative AI model analyzes the video data from within the factory, distinguishes between efficient and abnormal operations, and outputs improvement suggestions, allowing users to quickly identify problems in the manufacturing process and take measures.
[1353] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1354] Step 1:
[1355] Users use a dedicated device (tablet or smartphone) to record video of specific tasks within the factory and save it on the device.
[1356] Input: Video data of specific work being done in a factory
[1357] Output: Work video data saved on the device
[1358] Step 2:
[1359] The user uploads the work video data to the server through the terminal.
[1360] Input: Work video data saved on the device
[1361] Output: Work video data uploaded to the server
[1362] Step 3:
[1363] The server receives the uploaded work video data and extracts features using analysis software (e.g., OpenCV or TensorFlow).
[1364] Input: Work video data uploaded to the server
[1365] Output: Extracted features (worker movements, tool usage frequency, work time, etc.)
[1366] Step 4:
[1367] The server stores the extracted features as tags in a database.
[1368] Input: Extracted features
[1369] Output: Tag information stored in the database
[1370] Step 5:
[1371] The user uploads the actual manufacturing process video to the server via the terminal.
[1372] Input: Actual manufacturing process video data
[1373] Output: Manufacturing process video data uploaded to the server
[1374] Step 6:
[1375] The server analyzes the received manufacturing process video data and compares it with the aforementioned features.
[1376] Input: Manufacturing process video data uploaded to the server and tag information stored in the database
[1377] Output: Matching result (scenes that match the features)
[1378] Step 7:
[1379] Identify situations where the server is suspected to be abnormal or inefficient and organize them into a list.
[1380] Input: Matching result
[1381] Output: A list of suspected anomalies and inefficiencies
[1382] Step 8:
[1383] The user checks the list using a dedicated terminal and determines whether each scene is a predetermined work scene.
[1384] Input: List of suspected anomalies or inefficiencies
[1385] Output: Judgment result (whether it is a given work scene or not)
[1386] Step 9:
[1387] The judgment result is sent to the server, and the server updates the learning model.
[1388] Input: Judgment result
[1389] Output: Updated training model
[1390] Step 10:
[1391] The server generates improvement suggestions based on the analysis results and provides them to the user.
[1392] Input: Updated training model and analysis results
[1393] Output: Improvement suggestions
[1394] In this way, specific operations and data processing are carried out at each step, resulting in more efficient manufacturing processes and early detection of abnormalities.
[1395] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1396] System Overview
[1397] This invention is a system for learning and extracting specific scenes for the purpose of efficiently analyzing amateur sports game footage. This system has the ability to learn specific patterns and automatically extract specific scenes from game footage. Furthermore, by combining it with an emotion engine that recognizes user emotions and adjusts the analysis results based on those emotions, it is possible to improve user satisfaction.
[1398] Program processing
[1399] Learning Phase
[1400] 1. The user saves the video of the scene they want to extract (e.g., a set play video during practice) on their device and uploads it to the server via a dedicated application.
[1401] 2. The server receives the uploaded video and begins analysis.
[1402] 3. The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory. These features are stored as tags in a database.
[1403] Match video analysis phase
[1404] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[1405] 2. The server receives the match video and begins video analysis.
[1406] 3. The server analyzes the game footage frame by frame and compares it with pre-trained features. During this analysis process, scenes that match the features are picked out.
[1407] 4. The selected scenes are organized into a list and temporarily saved on the server.
[1408] Emotion Recognition Phase
[1409] 1. When a user uses a dedicated application to check the extracted scene clip, the device analyzes the user's facial expressions, voice, and gestures using an emotion engine.
[1410] 2. The emotion engine recognizes the user's emotions and sends that information to the server.
[1411] 3. The server receives the emotion engine data and dynamically adjusts the content and display method of the presented scene based on the user's emotions.
[1412] 4. The server records changes in the user's emotions and uses this information to improve the way scenes are presented and the accuracy of the analysis content in future sessions.
[1413] Verification Phase
[1414] 1. The user plays each clip and judges whether it is the intended scene. This judgment is made by a dedicated application, with a "correct" or "incorrect" result.
[1415] 2. The user's judgment results are sent to the server and recorded in a database. This data is used to update the system's learning model and improve the accuracy of analysis in future.
[1416] Specific examples
[1417] 1. The user takes a video of a "set play" practice (e.g., free kick practice) and saves the video on their device.
[1418] 2. The user uploads this video to the server using a dedicated application.
[1419] 3. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball trajectory and player positioning during a free kick are recorded as tags.
[1420] 4. The user takes video of the game and uploads it from their device to the server.
[1421] 5. The server receives the game footage and picks out scenes that match the learned features (e.g., free kick scenes).
[1422] 6. The server lists the scenes picked up and presents them to the user.
[1423] 7. When the user uses a dedicated application to check clips for each scene from the list, the device analyzes the user's facial expressions, voice, and gestures, and the emotion engine recognizes the user's emotions.
[1424] 8. The server dynamically adjusts how the clip is presented based on the user's emotions, for example, adjusting the playback speed of the clip if a positive emotion is detected, or revisiting a section of the clip if a negative emotion is detected.
[1425] 9. The user marks each clip as "true" or "false," and the result is sent to the server.
[1426] 10. The server records the judgment data and emotion data from the user, updates the system's learning model, and improves the accuracy of analysis from the next time onwards.
[1427] In this way, game video analysis is automated and dynamic scene adjustments based on the user's emotions are made possible, providing a more intuitive and satisfying video analysis experience. The system continues to learn and improves its analysis accuracy, resulting in an increasingly rich analysis experience.
[1428] The processing flow will be explained below.
[1429] Learning Phase
[1430] Step 1:
[1431] The video of the scene the user wants to extract (e.g., a set play video during practice) is saved on the device.
[1432] Step 2:
[1433] The user launches the dedicated application and selects the video of the scene they want to extract.
[1434] Step 3:
[1435] The terminal uploads the selected video to the server.
[1436] Step 4:
[1437] The server receives the uploaded video and begins analysis.
[1438] Step 5:
[1439] The server analyzes the video frame by frame and extracts features such as players' movements, positioning, and ball trajectory.
[1440] Step 6:
[1441] The server stores the extracted features as tags in a database.
[1442] Match video analysis phase
[1443] Step 1:
[1444] The user saves the game video on the device.
[1445] Step 2:
[1446] The user launches the dedicated application and selects the game footage.
[1447] Step 3:
[1448] The terminal uploads the selected game video to the server.
[1449] Step 4:
[1450] The server prepares to analyze the received game footage.
[1451] Step 5:
[1452] The server analyzes the game footage frame by frame and compares it with pre-learned features.
[1453] Step 6:
[1454] The server picks out scenes that match the matched features.
[1455] Step 7:
[1456] The server organizes the picked scenes into a list and temporarily saves them.
[1457] Emotion Recognition Phase
[1458] Step 1:
[1459] The user launches the dedicated application and checks the scene clips extracted from the server.
[1460] Step 2:
[1461] The device uses an emotion engine to analyze the user's emotional information, such as facial expressions, voice, and gestures.
[1462] Step 3:
[1463] The emotion engine recognizes the user's emotions and sends the information to the server.
[1464] Step 4:
[1465] The server receives the emotion engine data and dynamically adjusts the presentation method and content of the extracted scene based on the user's emotion.
[1466] Step 5:
[1467] The server records changes in the user's emotions and uses this information to improve the accuracy of scene presentations and analysis content in future sessions.
[1468] Verification Phase
[1469] Step 1:
[1470] The user plays each clip and determines whether it is the intended scene.
[1471] Step 2:
[1472] The user selects "true" or "false" for each clip in a dedicated application.
[1473] Step 3:
[1474] The user's decision result is sent to the server.
[1475] Step 4:
[1476] The server stores the judgment results and emotion recognition data in a database.
[1477] Step 5:
[1478] The server uses the stored data to update the learning model and improve the accuracy of analysis from the next time onwards.
[1479] Specific examples
[1480] Learning Phase
[1481] Step 1:
[1482] The user saves a "set play" practice video (e.g., free kick practice) on the device.
[1483] Step 2:
[1484] The user launches the dedicated application, selects a practice video, and uploads it to the server.
[1485] Step 3:
[1486] The device sends the video to the server.
[1487] Step 4:
[1488] The server receives the video and prepares it for analysis.
[1489] Step 5:
[1490] The server analyzes the video frame by frame and extracts features such as player movements, ball trajectory, and player positioning.
[1491] Step 6:
[1492] The server registers the extracted features in a database.
[1493] Match video analysis phase
[1494] Step 1:
[1495] The user saves the game video on the device.
[1496] Step 2:
[1497] The user launches the dedicated application, selects the game footage, and uploads it to the server.
[1498] Step 3:
[1499] The device sends the video to the server.
[1500] Step 4:
[1501] The server receives the match footage and prepares it for analysis.
[1502] Step 5:
[1503] The server analyzes the game footage frame by frame and compares it with the learned features.
[1504] Step 6:
[1505] The server picks out scenes that match the matched features (e.g., free kick scenes).
[1506] Step 7:
[1507] The server creates a list of the scenes it picks up and temporarily saves them.
[1508] Emotion Recognition Phase
[1509] Step 1:
[1510] The user starts the dedicated application and checks the clips of the scenes listed from the server.
[1511] Step 2:
[1512] The device analyzes the user's facial expressions, voice, gestures, etc. using an emotion engine.
[1513] Step 3:
[1514] The emotion engine recognizes the user's emotion and transmits the emotion information to the server.
[1515] Step 4:
[1516] The server receives the emotion engine data and dynamically adjusts how the clip is displayed based on the user's emotion.
[1517] Step 5:
[1518] The server records the user's emotional fluctuations and uses this information to improve the way scenes are presented and the accuracy of the analysis content from next time onwards.
[1519] Verification Phase
[1520] Step 1:
[1521] The user plays each clip and checks whether it is the intended scene.
[1522] Step 2:
[1523] The user selects "true" or "false" for each clip and sends the result to the server.
[1524] Step 3:
[1525] The server stores the judgment results and emotion recognition data in a database.
[1526] Step 4:
[1527] The server uses the stored data to update the learning model and improve analysis accuracy.
[1528] This series of processes automates the analysis of match footage and enables dynamic scene adjustments based on the user's emotions, providing a more intuitive and satisfying video analysis experience. The system continues to learn and improves its analysis accuracy, resulting in an increasingly rich analysis experience.
[1529] Example 2
[1530] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1531] When analyzing amateur sports game footage, it is difficult to efficiently and accurately extract specific scenes. Furthermore, to improve user satisfaction with the information obtained from video analysis, it is necessary to dynamically adjust the video based on the user's emotions. However, current systems lack the functionality to recognize the user's emotions and dynamically adjust the video presentation based on those emotions. As a result, users may not obtain the information they expect, potentially reducing the effectiveness of video analysis. A system that can solve these issues is needed.
[1532] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1533] In this invention, the server includes means for receiving video of scenes to be learned and extracting features from the video, means for receiving game video and extracting scenes from the game video that match the features, means for presenting the extracted scenes to a user and allowing the user to determine whether the scenes are predetermined scenes, means for recording the user's determination result and updating the learning model to improve extraction accuracy in the future, and means for recognizing the user's emotions and dynamically adjusting the content and presentation method of the extracted scenes based on the recognized emotions. This enables efficient and accurate scene extraction from game video and dynamic adjustment of the video in accordance with the user's emotions, thereby achieving video analysis with high user satisfaction.
[1534] Below are definitions of important words.
[1535] "Video of a scene to be learned" refers to video data that includes a specific scene that the user wants to extract.
[1536] "Features" refers to information extracted from video data that is necessary to identify specific scenes, such as player movements, positioning, and ball trajectory.
[1537] "Game footage" refers to video data that records the progress of an actual sports game.
[1538] An "emotion engine" refers to software or hardware that analyzes a user's facial expressions, voice, and gestures to recognize emotions.
[1539] A "learning model" refers to the algorithms and data structures that allow a system to continuously learn based on past data.
[1540] "Server" refers to a computer system that processes and stores video data and analysis data uploaded by users and performs various analyses.
[1541] "Terminal" refers to the device (e.g., smartphone, tablet, PC, etc.) that a user uses to store video and operate a dedicated application.
[1542] "Tag" refers to a label or metadata that is added to identify a specific feature in a video.
[1543] "Extraction" refers to extracting features or scenes that match predetermined conditions from video data.
[1544] "Dynamic adjustment" refers to changing the way the video is displayed in real time based on the user's emotions and other variables.
[1545] This allows the technical scope of the invention to be more clearly defined.
[1546] This invention is a system for efficiently analyzing amateur sports game footage, and has the ability to learn and extract specific scenes. This system learns specific patterns and automatically extracts specific scenes from game footage. Furthermore, by combining it with an emotion engine that recognizes user emotions and adjusts the analysis results based on those emotions, it is possible to improve user satisfaction.
[1547] The main hardware and software components for implementing this system are as follows:
[1548] 1. Device: A device used by a user to store video and operate dedicated applications. Examples include smartphones, tablets, and PCs.
[1549] 2. Server: A computer system that processes and stores video data and analysis data uploaded by users and performs various analyses. It is responsible for video analysis, database management, and execution of learning models.
[1550] 3. Emotion engine: Software or hardware that analyzes the user's facial expressions, voice, and gestures to recognize emotions.
[1551] The specific steps for implementing the system are as follows:
[1552] Learning Phase
[1553] 1. The user takes a video of the scene they want to extract and saves it on their device.
[1554] 2. The user uploads the video to the server through a dedicated application.
[1555] 3. The server receives the video and analyzes it frame by frame to extract features such as player movements, positioning, and ball trajectory.
[1556] Match video analysis phase
[1557] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[1558] 2. The server receives the game footage, analyzes each frame, and compares it with pre-trained features.
[1559] 3. The server picks up matching scenes, organizes them in a list format, and temporarily saves them.
[1560] Emotion Recognition Phase
[1561] 1. The user uses a dedicated application to view clips of the scenes selected here.
[1562] 2. The device sends the user's facial expressions, voice, and gestures to the emotion engine.
[1563] 3. The emotion engine generates emotion data and sends it to the server.
[1564] 4. The server dynamically adjusts the content and presentation of the clip based on the emotional data.
[1565] Verification Phase
[1566] 1. The user judges each clip as "true" or "false."
[1567] 2. The user's judgment result is sent to the server and recorded in the database.
[1568] 3. The server uses the judgment results to update the learning model and improve the accuracy of analysis from the next time onwards.
[1569] For example, if a user films and uploads a free-kick practice video, the server analyzes the ball's trajectory and player positions to select free-kick scenes from game footage. Then, when the user plays back the video, the system can dynamically adjust the playback speed and content of the clip based on the user's emotions.
[1570] Prompt Sentence Examples
[1571] 1. "Please explain the scenario where a user uploads a video of themselves practicing free kicks."
[1572] 2. "Please explain how to analyze frames from game footage."
[1573] 3. "Please explain how the emotion engine recognizes the user's emotions and how the server uses them."
[1574] This system enables efficient and accurate extraction of specific scenes, and also realizes dynamic adjustment of the video according to the user's emotions, thereby providing video analysis with high user satisfaction.
[1575] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1576] Learning Phase
[1577] Step 1:
[1578] The user takes a video of the scene they want to extract and saves it on their device.
[1579] Input: Video files captured by the camera.
[1580] Output: Video data saved on the device.
[1581] Step 2:
[1582] The user opens the dedicated application, selects the saved video file, and uploads it to the server.
[1583] Input: User specified video file.
[1584] Output: Video data received by the server.
[1585] Step 3:
[1586] The server starts a program to analyze the received video file.
[1587] Input: Uploaded video file.
[1588] Output: Video data divided into frames.
[1589] Step 4:
[1590] The server extracts features such as player movements, positioning, and ball trajectory for each frame.
[1591] Input: Video data divided into frames.
[1592] Output: Extracted feature data (player movements and positions, ball trajectory, etc.).
[1593] Step 5:
[1594] The server stores the extracted features as tags in a database.
[1595] Input: Extracted feature data.
[1596] Output: Tag information stored in the database.
[1597] Match video analysis phase
[1598] Step 1:
[1599] Users save game footage on their devices and upload it to the server via a dedicated application.
[1600] Input: User-specified game video file.
[1601] Output: Match video data received by the server.
[1602] Step 2:
[1603] The server receives the game footage and starts a program that analyzes it frame by frame.
[1604] Input: Uploaded match video file.
[1605] Output: Match video data divided into frames.
[1606] Step 3:
[1607] The server compares the features with those learned in advance and picks out scenes with matching features.
[1608] Input: Game video data divided into frames and pre-trained feature data.
[1609] Output: Frame number and time information of the matching scenes.
[1610] Step 4:
[1611] The server organizes the picked-up scenes in list format and temporarily saves them.
[1612] Input: Frame number and time information of the matching scene.
[1613] Output: List of scene information and its temporary storage.
[1614] Emotion Recognition Phase
[1615] Step 1:
[1616] The user can use a dedicated application to check clips of the scenes picked up here.
[1617] Input: Listed scene information.
[1618] Output: The clip of the scene being played.
[1619] Step 2:
[1620] The device transmits the user's facial expressions, voice, and gestures to the emotion engine.
[1621] Input: User facial, voice, and gesture data.
[1622] Output: Data sent to the emotion engine.
[1623] Step 3:
[1624] The emotion engine generates emotion data and sends it to the server.
[1625] Input: User facial, voice, and gesture data.
[1626] Output: The generated emotion data.
[1627] Step 4:
[1628] The server dynamically adjusts the content and presentation method of the clip based on the emotional data received.
[1629] Input: Emotion data and playing clip information.
[1630] Output: The playback speed and display of the adjusted clip.
[1631] Verification Phase
[1632] Step 1:
[1633] The user marks each clip as "true" or "false."
[1634] Input: The clip that was played.
[1635] Output: User's decision result.
[1636] Step 2:
[1637] The user's judgment result is sent to the server and recorded in a database.
[1638] Input: User's decision result.
[1639] Output: Verdict data recorded in a database.
[1640] Step 3:
[1641] The server uses the judgment results to update the learning model, improving the accuracy of analysis from the next time onwards.
[1642] Input: Recorded adjudication data.
[1643] Output: Updated training model.
[1644] (Application example 2)
[1645] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1646] When analyzing video of sports games, advertisements, etc., there is a need to efficiently extract specific scenes and dynamically adjust the display content based on the user's emotions. However, conventional technologies have had difficulty recognizing the user's emotions in real time and optimally displaying scenes and delivering advertisements in accordance with those emotions. The present invention aims to solve this problem and provide more intuitive and satisfying video analysis and advertisement delivery.
[1647] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1648] In this invention, the server includes means for receiving video of scenes to be learned and extracting features from the video, means for receiving game video and extracting scenes from the game video that match the features, and means for recognizing the user's emotions while watching and dynamically adjusting the content and display method of the extracted scenes based on the emotions, thereby enabling scene display and advertisement delivery in accordance with the user's emotions.
[1649] "Video of the scene to be learned" refers to video that the system uses to extract features in advance.
[1650] "Features" are data such as player movements, positioning, and ball trajectory that the system extracts from the video.
[1651] "Game footage" refers to footage of an actual sport or event being played.
[1652] "User" means any person or entity that uses the System.
[1653] "Means for adjusting based on emotions" is a mechanism that analyzes the user's emotions and dynamically changes the content and method of displaying the video based on the results.
[1654] "Dynamic adjustment" refers to changing the content and display method of the video in real time in response to changes in the user's emotions.
[1655] A "server" refers to a computer or network system that performs the main processing of a system.
[1656] System Overview
[1657] This invention is a system for efficiently analyzing sports game footage and advertising videos. It is characterized by recognizing the user's emotions while watching and dynamically adjusting the content and display method of the video based on those emotions. This system extracts features from the video of the scene being studied and analyzes the game footage and advertising videos in real time to provide the user with an optimal experience.
[1658] Program processing
[1659] Learning Phase
[1660] The server receives the video of the scene the user wants to extract (the learning video) and extracts features from the video. Specifically, it analyzes the video frame by frame and uses a generative AI model to extract data such as player movements, positioning, and ball trajectory. These features are then stored as tags in a database.
[1661] Match video analysis phase
[1662] Users save game footage or advertising footage on their devices and upload it to the server via a dedicated application. The server analyzes the received footage and compares it with pre-trained features to automatically select scenes that match. These scenes are organized into a list and temporarily saved.
[1663] Emotion Recognition Phase
[1664] When a user reviews an extracted scene using a dedicated application, the device analyzes the user's facial expressions, voice, and gestures using an emotion engine. The server receives the data sent from the emotion engine and dynamically adjusts how the scene is displayed based on the user's emotions. If a positive emotion is recognized, the playback speed of the clip is adjusted, and if a negative emotion is recognized, the clip is reviewed.
[1665] Hardware and software used
[1666] Hardware: High-performance servers, user devices (smartphones, tablets, PCs, etc.)
[1667] Software: Generative AI models (Keras, TensorFlow, OpenCV), emotion engine, database management system
[1668] Specific examples
[1669] For example, the present invention can be implemented as a system for delivering advertisements based on user emotions in the advertising industry. The server analyzes the advertisement video being viewed by the user in real time, recognizes the user's emotions while viewing, and dynamically selects the next advertisement to be delivered based on the emotions.
[1670] Examples of prompts:
[1671] "Please build a system that analyzes the advertisements that users watch in real time, recognizes their emotions from their facial expressions while watching, and dynamically selects the next advertisement to be delivered."
[1672] This makes it possible to deliver optimal video and advertisements according to the user's emotions, improving the viewing experience and maximizing the effectiveness of advertising.
[1673] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1674] Step 1:
[1675] Input: The user saves the video of the scene to be learned on the device.
[1676] How it works: The user uses a dedicated application to upload the video to be studied to the server.
[1677] Output: The learning video is saved on the server.
[1678] Step 2:
[1679] Input: The training video received by the server.
[1680] How it works: The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory, using generative AI models (Keras, TensorFlow).
[1681] Output: The extracted features are stored as tags in a database.
[1682] Step 3:
[1683] Input: The user saves game footage or advertising footage to the device.
[1684] How it works: Users use a dedicated application to upload game footage and advertising footage to the server.
[1685] Output: Match footage and advertising footage are saved on the server.
[1686] Step 4:
[1687] Input: Game footage and advertising footage received by the server, as well as pre-trained features.
[1688] How it works: The server analyzes the received video frame by frame and compares it with pre-trained features to pick out matching scenes. It then uses a generative AI model to confirm feature matches.
[1689] Output: Matching scenes are organized into a list and temporarily stored on the server.
[1690] Step 5:
[1691] Input: The user checks the extracted scene clips in a dedicated application.
[1692] How it works: While the user plays the clip, the device uses the camera and microphone to input the user's facial expressions, voice, and gestures into the emotion engine in real time.
[1693] Output: The emotion engine extracts the user's emotion data and sends it to the server.
[1694] Step 6:
[1695] Input: Emotion data received by the server, and a list of extracted scenes.
[1696] How it works: The server dynamically adjusts how the clip is presented based on the user's emotional data, for example adjusting the clip's playback speed if a positive emotion is detected, or revisiting a section of the clip if a negative emotion is detected.
[1697] Output: The adjusted scene is displayed to the user in real time.
[1698] Step 7:
[1699] Input: The user judges each clip as "true" or "false."
[1700] How it works: The user uses a dedicated application to send the results of their assessment of the clip to the server.
[1701] Output: The server records the user's judgment results in a database and updates the learning model to improve analysis accuracy in future.
[1702] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1703] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1704] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1705] [Fourth embodiment]
[1706] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1707] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1708] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1709] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1710] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1711] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1712] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1713] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1714] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1715] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1716] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1717] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1718] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1719] System Overview
[1720] This invention is a system for efficiently analyzing amateur sports game footage. This system has the function of learning specific scenes and automatically extracting specific scenes from game footage based on the learning data. This allows users to quickly check the intended scenes without any effort, greatly improving the efficiency of the analysis process.
[1721] Program processing
[1722] Learning Phase
[1723] 1. The user saves the video of the scene they want to extract (e.g., a set play video during practice) on their device and uploads it to the server via a dedicated application.
[1724] 2. The server receives the uploaded video and begins analysis.
[1725] 3. The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory. These features are stored as tags in a database.
[1726] Match video analysis phase
[1727] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[1728] 2. The server receives the match video and begins video analysis.
[1729] 3. The server analyzes the game footage frame by frame and compares it with pre-trained features. During this analysis process, scenes that match the features are picked out.
[1730] 4. The selected scenes are organized into a list and temporarily saved on the server.
[1731] Verification Phase
[1732] 1. The user checks the scene clips extracted from the server through a dedicated application.
[1733] 2. The user plays each clip and judges whether it is the intended scene. This judgment is made by a dedicated application, with a "correct" or "incorrect" result.
[1734] 3. The user's judgment results are sent to the server and recorded in a database. This data is used to update the system's learning model and improve analysis accuracy in future analyses.
[1735] Specific examples
[1736] 1. The user takes a video of a "set play" practice (e.g., free kick practice) and saves the video on their device.
[1737] 2. The user uploads this video to the server using a dedicated application.
[1738] 3. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball trajectory and player positioning during a free kick are recorded as tags.
[1739] 4. The user takes video of the game and uploads it from their device to the server.
[1740] 5. The server receives the game footage and picks out scenes that match the learned features (e.g., free kick scenes).
[1741] 6. The server lists the scenes picked up and presents them to the user.
[1742] 7. The user uses a dedicated application to check the clips of each scene from the list and determine whether they are the intended free kick scenes.
[1743] 8. The user judges whether the result is "correct" or "incorrect," and this information is sent to the server. The server uses this data to update the learning model and improve the accuracy of future analyses.
[1744] In this way, the analysis of game footage is automated, significantly reducing the user's workload. Furthermore, the system performs incremental learning, which continuously improves analysis accuracy and enables more accurate scene extraction.
[1745] The processing flow will be explained below.
[1746] Learning Phase
[1747] Step 1:
[1748] The video of the scene the user wants to extract (e.g., a set play video during practice) is saved on the device.
[1749] Step 2:
[1750] The user launches the dedicated application and selects the video of the scene they want to extract.
[1751] Step 3:
[1752] The terminal uploads the selected video to the server.
[1753] Step 4:
[1754] The server prepares to analyze the received video.
[1755] Step 5:
[1756] The server analyzes the video frame by frame and extracts features such as players' movements, positioning, and ball trajectory.
[1757] Step 6:
[1758] The server stores the extracted features as tags in a database.
[1759] Match video analysis phase
[1760] Step 1:
[1761] The user saves the game video on the device.
[1762] Step 2:
[1763] The user launches the dedicated application and selects the game footage.
[1764] Step 3:
[1765] The terminal uploads the selected game video to the server.
[1766] Step 4:
[1767] The server prepares to analyze the received game footage.
[1768] Step 5:
[1769] The server analyzes the game footage frame by frame and compares it with pre-learned features.
[1770] Step 6:
[1771] The server picks out scenes that match the matched features.
[1772] Step 7:
[1773] The server organizes the picked scenes into a list and temporarily saves them.
[1774] Verification Phase
[1775] Step 1:
[1776] The user launches the dedicated application and checks the scene clips extracted from the server.
[1777] Step 2:
[1778] The user plays each clip and determines whether it is the intended scene.
[1779] Step 3:
[1780] The user selects "true" or "false" for each clip in a dedicated application.
[1781] Step 4:
[1782] The user's decision result is sent to the server.
[1783] Step 5:
[1784] The server records the received judgment result in a database.
[1785] Step 6:
[1786] The server uses the recorded data to update the learning model, improving the accuracy of analysis from the next time onwards.
[1787] Specific examples
[1788] Learning Phase
[1789] Step 1:
[1790] The user saves a "set play" practice video (e.g., free kick practice) on the device.
[1791] Step 2:
[1792] The user launches the dedicated application, selects a practice video, and uploads it to the server.
[1793] Step 3:
[1794] The device sends the video to the server.
[1795] Step 4:
[1796] The server receives the video and prepares it for analysis.
[1797] Step 5:
[1798] The server analyzes the video frame by frame and extracts features such as player movements, ball trajectory, and player positioning.
[1799] Step 6:
[1800] The server registers the extracted features in a database.
[1801] Match video analysis phase
[1802] Step 1:
[1803] The user saves the game video on the device.
[1804] Step 2:
[1805] The user launches the dedicated application, selects the game footage, and uploads it to the server.
[1806] Step 3:
[1807] The device sends the video to the server.
[1808] Step 4:
[1809] The server receives the match footage and prepares it for analysis.
[1810] Step 5:
[1811] The server analyzes the game footage frame by frame and compares it with the learned features.
[1812] Step 6:
[1813] The server picks out scenes that match the matched features (e.g., free kick scenes).
[1814] Step 7:
[1815] The server creates a list of the scenes it picks up and temporarily saves them.
[1816] Verification Phase
[1817] Step 1:
[1818] The user starts the dedicated application and checks the clip of the scene notified by the server.
[1819] Step 2:
[1820] The user plays each clip and checks whether it is the intended scene.
[1821] Step 3:
[1822] The user uses a dedicated application to make a judgment ("correct" or "incorrect") about each clip.
[1823] Step 4:
[1824] The user's decision result is sent to the server.
[1825] Step 5:
[1826] The server stores the received judgment results in a database.
[1827] Step 6:
[1828] The server uses the stored data to update the analysis model and improve accuracy from the next time onwards.
[1829] This automates the analysis of game footage, allowing users to efficiently review the scenes they intended. The system continues to learn, improving its analysis accuracy and providing an increasingly rich analysis experience.
[1830] Example 1
[1831] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1832] In conventional amateur sports game video analysis systems, extracting specific scenes is a manual process that requires a great deal of time and effort. Furthermore, the accuracy of scene analysis is low, making it difficult to identify the desired scene. Furthermore, there are problems with systems that cannot handle changes in the video due to shooting at different distances or angles.
[1833] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1834] In this invention, the server includes: means for receiving video of a scene to be learned and extracting features from the video; means for receiving game video and extracting scenes from the game video that match the features; means for presenting the extracted scenes to a user and allowing the user to determine whether the scene is a specific scene; means for recording the user's determination and updating the learning model to improve extraction accuracy in future runs; means for analyzing the video frame by frame and extracting player movements, positioning, and ball trajectory; means for storing the uploaded video in a storage system; and means for selecting whether the determination result is "correct" or "incorrect." This not only enables automatic extraction of specific scenes with high accuracy, but also enables adaptation to different shooting conditions. Users can quickly and effortlessly confirm the intended scene, achieving efficient analysis work.
[1835] "Means for receiving video of a scene to be learned and extracting features from the video" refers to a device or method in which a server takes in video data containing a specific scene provided by a user, and analyzes and extracts identifiable elements from the video, such as player movements, positioning, and ball trajectory.
[1836] The "means for receiving game footage and extracting scenes from the game footage that match the features" is a mechanism by which the server receives game footage provided by the user, compares it with features learned in advance, and automatically extracts matching parts.
[1837] The "means of presenting the extracted scene to the user and allowing the user to determine whether the scene is a specified scene" refers to a method in which the server displays the video clip extracted as the analysis result to the user through a user interface, allowing the user to check and determine whether the clip is the scene they want.
[1838] "Means for recording the user's judgment results and updating the learning model to improve extraction accuracy in the future" refers to the process in which the server stores the judgment results ("correct" or "incorrect") provided by the user in a database, and retrains and updates the machine learning model based on them to improve the accuracy of subsequent analysis.
[1839] "Means of analyzing video frame by frame to extract player movements, positioning, and ball trajectory" refers to a technology that divides video data into fixed time intervals and analyzes each frame in detail to extract important information such as the movements and positioning of players and the ball.
[1840] The "means for storing uploaded video in a storage system" refers to online storage or a database for temporarily or long-term storage of video data sent by a user to a server.
[1841] The "means for selecting whether the judgment result is 'correct' or 'incorrect'" is a user interface that allows the user to select 'correct' or 'incorrect' as to whether the content of a presented video clip matches a specified scene.
[1842] System Overview
[1843] This invention is a system for efficiently analyzing amateur sports game footage. This system has the function of learning specific scenes and automatically extracting specific scenes from game footage based on the learning data. This allows users to quickly check the intended scenes without any effort, greatly improving the efficiency of the analysis process.
[1844] Hardware and software used
[1845] The hardware used to implement this system includes general servers, terminals, and devices that provide the user interface. The server requires a high-performance CPU and a large amount of storage, and it is recommended to use a cloud service (e.g., AWS or Google Cloud). The terminal can be a smartphone, tablet, or PC.
[1846] The software uses OpenCV for video analysis, TensorFlow or PyTorch for machine learning, and MySQL for database management. The dedicated application allows users to upload recorded video to the server and provides an interface for checking the analysis results.
[1847] Program processing
[1848] Learning Phase
[1849] The user saves the video of the scene they want to extract (for example, a set play during practice) on their device. This video is then uploaded to the server via a dedicated application. The server analyzes the received video frame by frame, extracting features such as player movements, positioning, and ball trajectory, and stores these features as tags in a database.
[1850] Match video analysis phase
[1851] The user saves the game footage on their device and uploads it to the server via a dedicated application. The server receives the game footage and begins video analysis. The game footage is analyzed frame by frame and compared with pre-trained feature vectors. During this analysis process, scenes that match the feature vectors are picked out, organized into a list, and temporarily stored on the server.
[1852] Verification Phase
[1853] The user checks the scene clips extracted from the server through a dedicated application. Each clip is played and judged to be the intended scene. The judgement is made in the dedicated application as "correct" or "incorrect." The user's judgement is sent to the server and recorded in a database. This data is used to update the system's learning model and improve the accuracy of analysis from the next time onwards.
[1854] Specific examples
[1855] The user films a "set play" practice video (e.g., free kick practice) and saves the video on their device. The user then uploads this video to a server using a dedicated application. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball's trajectory and player positioning during a free kick are recorded as tags. The user films a game during the match and uploads the video from their device to the server. The server receives the game video and selects scenes (e.g., free kick scenes) that match the learned features. The server then lists the selected scenes and presents them to the user. The user then uses a dedicated application to review each scene clip from the list and determine whether it is the intended free kick scene. The user then judges whether it is "correct" or "incorrect," and this information is sent to the server. The server uses this data to update the learning model, improving the accuracy of future analysis.
[1856] Prompt Sentence Examples
[1857] "Please explain how to film amateur soccer free kick practice scenes, save them on a device, and upload the footage to a server using a dedicated application. Also, please explain the steps for how the server analyzes and learns from the uploaded footage."
[1858] In this way, the analysis of game footage is automated, significantly reducing the user's workload. Furthermore, the system performs incremental learning, which continuously improves analysis accuracy and enables more accurate scene extraction.
[1859] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1860] Step 1:
[1861] The video of the scene that the user wants to extract (e.g., a set play video during practice) is recorded and saved on the device. The input is a video file containing a specific scene created by the user, and the output is a video file saved on the device. The user uses a smartphone or camera to record the set practice scene and saves the data on the device.
[1862] Step 2:
[1863] The user launches the dedicated application and selects the recorded video file within the application. Then, they click the "Upload" button in the application to upload the video file to the server. The input is the video file saved on the device, and the output is the video data uploaded to the server.
[1864] Step 3:
[1865] The server receives the uploaded video file. It stores the received video file in a storage system (e.g., AWS S3). The input is the uploaded video file, and the output is the stored video data.
[1866] Step 4:
[1867] The server analyzes the video frame by frame. Video processing software such as OpenCV is used for video analysis. Through the analysis, features such as player movements, positioning, and ball trajectory are extracted and stored as tags in a database. The input is the stored video data, and the output is tag data of the extracted features.
[1868] Step 5:
[1869] The user saves the game video on the device. The entire game video is recorded and the data is saved on the device. The input is the game video, and the output is the game video saved on the device. The user also records the game video and saves it on the device.
[1870] Step 6:
[1871] The user uploads game footage saved on their device to the server using a dedicated application. The input is the game footage file saved on the device, and the output is the game footage data uploaded to the server.
[1872] Step 7:
[1873] The server receives and stores the game footage. It analyzes the received footage frame by frame and compares it with pre-trained features. Computer vision technology and machine learning models (e.g., TensorFlow) are used for the matching process. The input is the stored game footage data, and the output is the extraction of scenes that match the features.
[1874] Step 8:
[1875] The server compiles a list of matched scenes and temporarily stores it. The input is the frame data of the matched scenes, and the output is a scene list. These scenes are organized for presentation to the user.
[1876] Step 9:
[1877] The user launches a dedicated application and checks the scene list provided by the server. The user selects the "Check Scene" option within the application and plays a clip of the presented scene. The input is the list provided by the server, and the output is a clip of the scene selected by the user.
[1878] Step 10:
[1879] The user plays each clip and judges whether it is the intended scene as "correct" or "incorrect." This judgment is made on the interface of a dedicated application. The input is the played clip video, and the output is the judgment result, "correct" or "incorrect."
[1880] Step 11:
[1881] The user's judgment results are sent to the server via a dedicated application. The server stores the judgment results in a database and uses this data to retrain and update the machine learning model. The input is the user's judgment results, and the output is the updated learning model. This improves the accuracy of analysis from the next time onwards.
[1882] (Application example 1)
[1883] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1884] This invention relates to a system that efficiently analyzes the manufacturing processes of automated robots and workers used in factories, automatically extracts and evaluates specific work scenes, and improves manufacturing efficiency and detects anomalies. In current manufacturing processes, it is difficult to quickly detect and analyze inefficient operations or abnormalities, which can result in adverse effects on product quality and productivity. To solve this problem, technology is needed that can automatically monitor and analyze manufacturing processes and extract specific work scenes.
[1885] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1886] In this invention, the server includes means for receiving video of a work scene to be learned and extracting features from the video, means for receiving manufacturing process video and extracting scenes from the manufacturing process video that match the features, means for presenting the extracted work scene to a user and allowing the user to determine whether the scene is a predetermined work scene, means for recording the user's determination result and updating the learning model to improve extraction accuracy from the next time onwards, means for analyzing scenes suspected of abnormalities or reduced efficiency and organizing them into a list, and means for generating improvement suggestions based on the analysis results, thereby enabling efficiency improvement in the manufacturing process and early detection of abnormalities.
[1887] "Video of a work scene to be learned" is video data that records the process of a specific work being done in a factory.
[1888] "Features" are data extracted from video data, such as the worker's movements, frequency of tool use, and work time.
[1889] "Manufacturing process video" is video data that records the actual manufacturing process.
[1890] A "matching scene" is a part of the manufacturing process video that has features similar to the learned features.
[1891] "User" refers to a person in a factory who uses the system to analyze the manufacturing process.
[1892] A "predetermined work scene" is a video scene showing the process of a specific work that has been determined in advance.
[1893] The "determination result" is the result of the user's determination as to whether or not the scene is a predetermined work scene.
[1894] A "learning model" is a data processing algorithm that learns features based on video data and improves analysis accuracy.
[1895] "Scenes where abnormalities or reduced efficiency are suspected" are scenes where abnormal operations or reduced work efficiency are observed compared to the normal manufacturing process.
[1896] A "list" is data that organizes and lists situations where abnormalities or reduced efficiency are suspected.
[1897] "Improvement proposals" are specific action plans proposed based on the analysis results to improve the efficiency of manufacturing processes and correct abnormalities.
[1898] The system for implementing this invention starts by recording the manufacturing process of automated robots and workers used in a factory with a camera and uploading the video data to a server. The server then uses dedicated software for analyzing the video data (e.g., OpenCV or TensorFlow) to extract features such as the worker's movements, frequency of tool use, and work time for each frame and stores them in a database.
[1899] Next, actual manufacturing process footage is similarly uploaded to the server and compared with the pre-trained feature values. At this time, the server picks out scenes that are suspected of being abnormal or inefficient and organizes them into a list. The user checks this list using a dedicated device (tablet or smartphone) and determines whether each scene is a specified work scene. The results of this determination are sent to the server, and the learning model is updated, improving extraction accuracy from the next time onwards.
[1900] Furthermore, the server generates improvement proposals based on the analysis results and provides them to users. This enables the efficiency of the manufacturing process and the early detection of abnormalities. In addition, the system has the advantage of being able to continuously improve its analysis accuracy because it learns sequentially.
[1901] Hardware used
[1902] Camera (installed inside the factory)
[1903] Factory Robots
[1904] Dedicated tablet / smartphone
[1905] Servers (e.g., high-performance computing resources on Amazon Web Services (AWS) or Google Cloud Platform (GCP))
[1906] Software used
[1907] Analysis software (e.g., OpenCV (image processing library), TensorFlow (machine learning library))
[1908] Database management system (e.g. MySQL)
[1909] Specific examples
[1910] For example, to analyze a video of a scene in an assembly room in a factory where a robot arm is performing a specific movement and detect a specific movement pattern, the following prompt sentence would be used.
[1911] Prompt Sentence Examples
[1912] "During factory assembly work, identify instances where the robot arm's movement slows down or stops, and compare it with manual work."
[1913] Based on this prompt, the generative AI model analyzes the video data from within the factory, distinguishes between efficient and abnormal operations, and outputs improvement suggestions, allowing users to quickly identify problems in the manufacturing process and take measures.
[1914] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1915] Step 1:
[1916] Users use a dedicated device (tablet or smartphone) to record video of specific tasks within the factory and save it on the device.
[1917] Input: Video data of specific work being done in a factory
[1918] Output: Work video data saved on the device
[1919] Step 2:
[1920] The user uploads the work video data to the server through the terminal.
[1921] Input: Work video data saved on the device
[1922] Output: Work video data uploaded to the server
[1923] Step 3:
[1924] The server receives the uploaded work video data and extracts features using analysis software (e.g., OpenCV or TensorFlow).
[1925] Input: Work video data uploaded to the server
[1926] Output: Extracted features (worker movements, tool usage frequency, work time, etc.)
[1927] Step 4:
[1928] The server stores the extracted features as tags in a database.
[1929] Input: Extracted features
[1930] Output: Tag information stored in the database
[1931] Step 5:
[1932] The user uploads the actual manufacturing process video to the server via the terminal.
[1933] Input: Actual manufacturing process video data
[1934] Output: Manufacturing process video data uploaded to the server
[1935] Step 6:
[1936] The server analyzes the received manufacturing process video data and compares it with the aforementioned features.
[1937] Input: Manufacturing process video data uploaded to the server and tag information stored in the database
[1938] Output: Matching result (scenes that match the features)
[1939] Step 7:
[1940] Identify situations where the server is suspected to be abnormal or inefficient and organize them into a list.
[1941] Input: Matching result
[1942] Output: A list of suspected anomalies and inefficiencies
[1943] Step 8:
[1944] The user checks the list using a dedicated terminal and determines whether each scene is a predetermined work scene.
[1945] Input: List of suspected anomalies or inefficiencies
[1946] Output: Judgment result (whether it is a given work scene or not)
[1947] Step 9:
[1948] The judgment result is sent to the server, and the server updates the learning model.
[1949] Input: Judgment result
[1950] Output: Updated training model
[1951] Step 10:
[1952] The server generates improvement suggestions based on the analysis results and provides them to the user.
[1953] Input: Updated training model and analysis results
[1954] Output: Improvement suggestions
[1955] In this way, specific operations and data processing are carried out at each step, resulting in more efficient manufacturing processes and early detection of abnormalities.
[1956] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1957] System Overview
[1958] This invention is a system for learning and extracting specific scenes for the purpose of efficiently analyzing amateur sports game footage. This system has the ability to learn specific patterns and automatically extract specific scenes from game footage. Furthermore, by combining it with an emotion engine that recognizes user emotions and adjusts the analysis results based on those emotions, it is possible to improve user satisfaction.
[1959] Program processing
[1960] Learning Phase
[1961] 1. The user saves the video of the scene they want to extract (e.g., a set play video during practice) on their device and uploads it to the server via a dedicated application.
[1962] 2. The server receives the uploaded video and begins analysis.
[1963] 3. The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory. These features are stored as tags in a database.
[1964] Match video analysis phase
[1965] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[1966] 2. The server receives the match video and begins video analysis.
[1967] 3. The server analyzes the game footage frame by frame and compares it with pre-trained features. During this analysis process, scenes that match the features are picked out.
[1968] 4. The selected scenes are organized into a list and temporarily saved on the server.
[1969] Emotion Recognition Phase
[1970] 1. When a user uses a dedicated application to check the extracted scene clip, the device analyzes the user's facial expressions, voice, and gestures using an emotion engine.
[1971] 2. The emotion engine recognizes the user's emotions and sends that information to the server.
[1972] 3. The server receives the emotion engine data and dynamically adjusts the content and display method of the presented scene based on the user's emotions.
[1973] 4. The server records changes in the user's emotions and uses this information to improve the way scenes are presented and the accuracy of the analysis content in future sessions.
[1974] Verification Phase
[1975] 1. The user plays each clip and judges whether it is the intended scene. This judgment is made by a dedicated application, with a "correct" or "incorrect" result.
[1976] 2. The user's judgment results are sent to the server and recorded in a database. This data is used to update the system's learning model and improve the accuracy of analysis in future.
[1977] Specific examples
[1978] 1. The user takes a video of a "set play" practice (e.g., free kick practice) and saves the video on their device.
[1979] 2. The user uploads this video to the server using a dedicated application.
[1980] 3. The server receives the video and extracts features such as player movements, positioning, and ball trajectory. For example, the ball trajectory and player positioning during a free kick are recorded as tags.
[1981] 4. The user takes video of the game and uploads it from their device to the server.
[1982] 5. The server receives the game footage and picks out scenes that match the learned features (e.g., free kick scenes).
[1983] 6. The server lists the scenes picked up and presents them to the user.
[1984] 7. When the user uses a dedicated application to check clips for each scene from the list, the device analyzes the user's facial expressions, voice, and gestures, and the emotion engine recognizes the user's emotions.
[1985] 8. The server dynamically adjusts how the clip is presented based on the user's emotions, for example, adjusting the playback speed of the clip if a positive emotion is detected, or revisiting a section of the clip if a negative emotion is detected.
[1986] 9. The user marks each clip as "true" or "false," and the result is sent to the server.
[1987] 10. The server records the judgment data and emotion data from the user, updates the system's learning model, and improves the accuracy of analysis from the next time onwards.
[1988] In this way, game video analysis is automated and dynamic scene adjustments based on the user's emotions are made possible, providing a more intuitive and satisfying video analysis experience. The system continues to learn and improves its analysis accuracy, resulting in an increasingly rich analysis experience.
[1989] The processing flow will be explained below.
[1990] Learning Phase
[1991] Step 1:
[1992] The video of the scene the user wants to extract (e.g., a set play video during practice) is saved on the device.
[1993] Step 2:
[1994] The user launches the dedicated application and selects the video of the scene they want to extract.
[1995] Step 3:
[1996] The terminal uploads the selected video to the server.
[1997] Step 4:
[1998] The server receives the uploaded video and begins analysis.
[1999] Step 5:
[2000] The server analyzes the video frame by frame and extracts features such as players' movements, positioning, and ball trajectory.
[2001] Step 6:
[2002] The server stores the extracted features as tags in a database.
[2003] Match video analysis phase
[2004] Step 1:
[2005] The user saves the game video on the device.
[2006] Step 2:
[2007] The user launches the dedicated application and selects the game footage.
[2008] Step 3:
[2009] The terminal uploads the selected game video to the server.
[2010] Step 4:
[2011] The server prepares to analyze the received game footage.
[2012] Step 5:
[2013] The server analyzes the game footage frame by frame and compares it with pre-learned features.
[2014] Step 6:
[2015] The server picks out scenes that match the matched features.
[2016] Step 7:
[2017] The server organizes the picked scenes into a list and temporarily saves them.
[2018] Emotion Recognition Phase
[2019] Step 1:
[2020] The user launches the dedicated application and checks the scene clips extracted from the server.
[2021] Step 2:
[2022] The device uses an emotion engine to analyze the user's emotional information, such as facial expressions, voice, and gestures.
[2023] Step 3:
[2024] The emotion engine recognizes the user's emotions and sends the information to the server.
[2025] Step 4:
[2026] The server receives the emotion engine data and dynamically adjusts the presentation method and content of the extracted scene based on the user's emotion.
[2027] Step 5:
[2028] The server records changes in the user's emotions and uses this information to improve the accuracy of scene presentations and analysis content in future sessions.
[2029] Verification Phase
[2030] Step 1:
[2031] The user plays each clip and determines whether it is the intended scene.
[2032] Step 2:
[2033] The user selects "true" or "false" for each clip in a dedicated application.
[2034] Step 3:
[2035] The user's decision result is sent to the server.
[2036] Step 4:
[2037] The server stores the judgment results and emotion recognition data in a database.
[2038] Step 5:
[2039] The server uses the stored data to update the learning model and improve the accuracy of analysis from the next time onwards.
[2040] Specific examples
[2041] Learning Phase
[2042] Step 1:
[2043] The user saves a "set play" practice video (e.g., free kick practice) on the device.
[2044] Step 2:
[2045] The user launches the dedicated application, selects a practice video, and uploads it to the server.
[2046] Step 3:
[2047] The device sends the video to the server.
[2048] Step 4:
[2049] The server receives the video and prepares it for analysis.
[2050] Step 5:
[2051] The server analyzes the video frame by frame and extracts features such as player movements, ball trajectory, and player positioning.
[2052] Step 6:
[2053] The server registers the extracted features in a database.
[2054] Match video analysis phase
[2055] Step 1:
[2056] The user saves the game video on the device.
[2057] Step 2:
[2058] The user launches the dedicated application, selects the game footage, and uploads it to the server.
[2059] Step 3:
[2060] The device sends the video to the server.
[2061] Step 4:
[2062] The server receives the match footage and prepares it for analysis.
[2063] Step 5:
[2064] The server analyzes the game footage frame by frame and compares it with the learned features.
[2065] Step 6:
[2066] The server picks out scenes that match the matched features (e.g., free kick scenes).
[2067] Step 7:
[2068] The server creates a list of the scenes it picks up and temporarily saves them.
[2069] Emotion Recognition Phase
[2070] Step 1:
[2071] The user starts the dedicated application and checks the clips of the scenes listed from the server.
[2072] Step 2:
[2073] The device analyzes the user's facial expressions, voice, gestures, etc. using an emotion engine.
[2074] Step 3:
[2075] The emotion engine recognizes the user's emotion and transmits the emotion information to the server.
[2076] Step 4:
[2077] The server receives the emotion engine data and dynamically adjusts how the clip is displayed based on the user's emotion.
[2078] Step 5:
[2079] The server records the user's emotional fluctuations and uses this information to improve the way scenes are presented and the accuracy of the analysis content from next time onwards.
[2080] Verification Phase
[2081] Step 1:
[2082] The user plays each clip and checks whether it is the intended scene.
[2083] Step 2:
[2084] The user selects "true" or "false" for each clip and sends the result to the server.
[2085] Step 3:
[2086] The server stores the judgment results and emotion recognition data in a database.
[2087] Step 4:
[2088] The server uses the stored data to update the learning model and improve analysis accuracy.
[2089] This series of processes automates the analysis of match footage and enables dynamic scene adjustments based on the user's emotions, providing a more intuitive and satisfying video analysis experience. The system continues to learn and improves its analysis accuracy, resulting in an increasingly rich analysis experience.
[2090] Example 2
[2091] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2092] When analyzing amateur sports game footage, it is difficult to efficiently and accurately extract specific scenes. Furthermore, to improve user satisfaction with the information obtained from video analysis, it is necessary to dynamically adjust the video based on the user's emotions. However, current systems lack the functionality to recognize the user's emotions and dynamically adjust the video presentation based on those emotions. As a result, users may not obtain the information they expect, potentially reducing the effectiveness of video analysis. A system that can solve these issues is needed.
[2093] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2094] In this invention, the server includes means for receiving video of scenes to be learned and extracting features from the video, means for receiving game video and extracting scenes from the game video that match the features, means for presenting the extracted scenes to a user and allowing the user to determine whether the scenes are predetermined scenes, means for recording the user's determination result and updating the learning model to improve extraction accuracy in the future, and means for recognizing the user's emotions and dynamically adjusting the content and presentation method of the extracted scenes based on the recognized emotions. This enables efficient and accurate scene extraction from game video and dynamic adjustment of the video in accordance with the user's emotions, thereby achieving video analysis with high user satisfaction.
[2095] Below are definitions of important words.
[2096] "Video of a scene to be learned" refers to video data that includes a specific scene that the user wants to extract.
[2097] "Features" refers to information extracted from video data that is necessary to identify specific scenes, such as player movements, positioning, and ball trajectory.
[2098] "Game footage" refers to video data that records the progress of an actual sports game.
[2099] An "emotion engine" refers to software or hardware that analyzes a user's facial expressions, voice, and gestures to recognize emotions.
[2100] A "learning model" refers to the algorithms and data structures that allow a system to continuously learn based on past data.
[2101] "Server" refers to a computer system that processes and stores video data and analysis data uploaded by users and performs various analyses.
[2102] "Terminal" refers to the device (e.g., smartphone, tablet, PC, etc.) that a user uses to store video and operate a dedicated application.
[2103] "Tag" refers to a label or metadata that is added to identify a specific feature in a video.
[2104] "Extraction" refers to extracting features or scenes that match predetermined conditions from video data.
[2105] "Dynamic adjustment" refers to changing the way the video is displayed in real time based on the user's emotions and other variables.
[2106] This allows the technical scope of the invention to be more clearly defined.
[2107] This invention is a system for efficiently analyzing amateur sports game footage, and has the ability to learn and extract specific scenes. This system learns specific patterns and automatically extracts specific scenes from game footage. Furthermore, by combining it with an emotion engine that recognizes user emotions and adjusts the analysis results based on those emotions, it is possible to improve user satisfaction.
[2108] The main hardware and software components for implementing this system are as follows:
[2109] 1. Device: A device used by a user to store video and operate dedicated applications. Examples include smartphones, tablets, and PCs.
[2110] 2. Server: A computer system that processes and stores video data and analysis data uploaded by users and performs various analyses. It is responsible for video analysis, database management, and execution of learning models.
[2111] 3. Emotion engine: Software or hardware that analyzes the user's facial expressions, voice, and gestures to recognize emotions.
[2112] The specific steps for implementing the system are as follows:
[2113] Learning Phase
[2114] 1. The user takes a video of the scene they want to extract and saves it on their device.
[2115] 2. The user uploads the video to the server through a dedicated application.
[2116] 3. The server receives the video and analyzes it frame by frame to extract features such as player movements, positioning, and ball trajectory.
[2117] Match video analysis phase
[2118] 1. The user saves the game footage on their device and uploads it to the server via a dedicated application.
[2119] 2. The server receives the game footage, analyzes each frame, and compares it with pre-trained features.
[2120] 3. The server picks up matching scenes, organizes them in a list format, and temporarily saves them.
[2121] Emotion Recognition Phase
[2122] 1. The user uses a dedicated application to view clips of the scenes selected here.
[2123] 2. The device sends the user's facial expressions, voice, and gestures to the emotion engine.
[2124] 3. The emotion engine generates emotion data and sends it to the server.
[2125] 4. The server dynamically adjusts the content and presentation of the clip based on the emotional data.
[2126] Verification Phase
[2127] 1. The user judges each clip as "true" or "false."
[2128] 2. The user's judgment result is sent to the server and recorded in the database.
[2129] 3. The server uses the judgment results to update the learning model and improve the accuracy of analysis from the next time onwards.
[2130] For example, if a user films and uploads a free-kick practice video, the server analyzes the ball's trajectory and player positions to select free-kick scenes from game footage. Then, when the user plays back the video, the system can dynamically adjust the playback speed and content of the clip based on the user's emotions.
[2131] Prompt Sentence Examples
[2132] 1. "Please explain the scenario where a user uploads a video of themselves practicing free kicks."
[2133] 2. "Please explain how to analyze frames from game footage."
[2134] 3. "Please explain how the emotion engine recognizes the user's emotions and how the server uses them."
[2135] This system enables efficient and accurate extraction of specific scenes, and also realizes dynamic adjustment of the video according to the user's emotions, thereby providing video analysis with high user satisfaction.
[2136] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2137] Learning Phase
[2138] Step 1:
[2139] The user takes a video of the scene they want to extract and saves it on their device.
[2140] Input: Video files captured by the camera.
[2141] Output: Video data saved on the device.
[2142] Step 2:
[2143] The user opens the dedicated application, selects the saved video file, and uploads it to the server.
[2144] Input: User specified video file.
[2145] Output: Video data received by the server.
[2146] Step 3:
[2147] The server starts a program to analyze the received video file.
[2148] Input: Uploaded video file.
[2149] Output: Video data divided into frames.
[2150] Step 4:
[2151] The server extracts features such as player movements, positioning, and ball trajectory for each frame.
[2152] Input: Video data divided into frames.
[2153] Output: Extracted feature data (player movements and positions, ball trajectory, etc.).
[2154] Step 5:
[2155] The server stores the extracted features as tags in a database.
[2156] Input: Extracted feature data.
[2157] Output: Tag information stored in the database.
[2158] Match video analysis phase
[2159] Step 1:
[2160] Users save game footage on their devices and upload it to the server via a dedicated application.
[2161] Input: User-specified game video file.
[2162] Output: Match video data received by the server.
[2163] Step 2:
[2164] The server receives the game footage and starts a program that analyzes it frame by frame.
[2165] Input: Uploaded match video file.
[2166] Output: Match video data divided into frames.
[2167] Step 3:
[2168] The server compares the features with those learned in advance and picks out scenes with matching features.
[2169] Input: Game video data divided into frames and pre-trained feature data.
[2170] Output: Frame number and time information of the matching scenes.
[2171] Step 4:
[2172] The server organizes the picked-up scenes in list format and temporarily saves them.
[2173] Input: Frame number and time information of the matching scene.
[2174] Output: List of scene information and its temporary storage.
[2175] Emotion Recognition Phase
[2176] Step 1:
[2177] The user can use a dedicated application to check clips of the scenes picked up here.
[2178] Input: Listed scene information.
[2179] Output: The clip of the scene being played.
[2180] Step 2:
[2181] The device transmits the user's facial expressions, voice, and gestures to the emotion engine.
[2182] Input: User facial, voice, and gesture data.
[2183] Output: Data sent to the emotion engine.
[2184] Step 3:
[2185] The emotion engine generates emotion data and sends it to the server.
[2186] Input: User facial, voice, and gesture data.
[2187] Output: The generated emotion data.
[2188] Step 4:
[2189] The server dynamically adjusts the content and presentation method of the clip based on the emotional data received.
[2190] Input: Emotion data and playing clip information.
[2191] Output: The playback speed and display of the adjusted clip.
[2192] Verification Phase
[2193] Step 1:
[2194] The user marks each clip as "true" or "false."
[2195] Input: The clip that was played.
[2196] Output: User's decision result.
[2197] Step 2:
[2198] The user's judgment result is sent to the server and recorded in a database.
[2199] Input: User's decision result.
[2200] Output: Verdict data recorded in a database.
[2201] Step 3:
[2202] The server uses the judgment results to update the learning model, improving the accuracy of analysis from the next time onwards.
[2203] Input: Recorded adjudication data.
[2204] Output: Updated training model.
[2205] (Application example 2)
[2206] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2207] When analyzing video of sports games, advertisements, etc., there is a need to efficiently extract specific scenes and dynamically adjust the display content based on the user's emotions. However, conventional technologies have had difficulty recognizing the user's emotions in real time and optimally displaying scenes and delivering advertisements in accordance with those emotions. The present invention aims to solve this problem and provide more intuitive and satisfying video analysis and advertisement delivery.
[2208] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2209] In this invention, the server includes means for receiving video of scenes to be learned and extracting features from the video, means for receiving game video and extracting scenes from the game video that match the features, and means for recognizing the user's emotions while watching and dynamically adjusting the content and display method of the extracted scenes based on the emotions, thereby enabling scene display and advertisement delivery in accordance with the user's emotions.
[2210] "Video of the scene to be learned" refers to video that the system uses to extract features in advance.
[2211] "Features" are data such as player movements, positioning, and ball trajectory that the system extracts from the video.
[2212] "Game footage" refers to footage of an actual sport or event being played.
[2213] "User" means any person or entity that uses the System.
[2214] "Means for adjusting based on emotions" is a mechanism that analyzes the user's emotions and dynamically changes the content and method of displaying the video based on the results.
[2215] "Dynamic adjustment" refers to changing the content and display method of the video in real time in response to changes in the user's emotions.
[2216] A "server" refers to a computer or network system that performs the main processing of a system.
[2217] System Overview
[2218] This invention is a system for efficiently analyzing sports game footage and advertising videos. It is characterized by recognizing the user's emotions while watching and dynamically adjusting the content and display method of the video based on those emotions. This system extracts features from the video of the scene being studied and analyzes the game footage and advertising videos in real time to provide the user with an optimal experience.
[2219] Program processing
[2220] Learning Phase
[2221] The server receives the video of the scene the user wants to extract (the learning video) and extracts features from the video. Specifically, it analyzes the video frame by frame and uses a generative AI model to extract data such as player movements, positioning, and ball trajectory. These features are then stored as tags in a database.
[2222] Match video analysis phase
[2223] Users save game footage or advertising footage on their devices and upload it to the server via a dedicated application. The server analyzes the received footage and compares it with pre-trained features to automatically select scenes that match. These scenes are organized into a list and temporarily saved.
[2224] Emotion Recognition Phase
[2225] When a user reviews an extracted scene using a dedicated application, the device analyzes the user's facial expressions, voice, and gestures using an emotion engine. The server receives the data sent from the emotion engine and dynamically adjusts how the scene is displayed based on the user's emotions. If a positive emotion is recognized, the playback speed of the clip is adjusted, and if a negative emotion is recognized, the clip is reviewed.
[2226] Hardware and software used
[2227] Hardware: High-performance servers, user devices (smartphones, tablets, PCs, etc.)
[2228] Software: Generative AI models (Keras, TensorFlow, OpenCV), emotion engine, database management system
[2229] Specific examples
[2230] For example, the present invention can be implemented as a system for delivering advertisements based on user emotions in the advertising industry. The server analyzes the advertisement video being viewed by the user in real time, recognizes the user's emotions while viewing, and dynamically selects the next advertisement to be delivered based on the emotions.
[2231] Examples of prompts:
[2232] "Please build a system that analyzes the advertisements that users watch in real time, recognizes their emotions from their facial expressions while watching, and dynamically selects the next advertisement to be delivered."
[2233] This makes it possible to deliver optimal video and advertisements according to the user's emotions, improving the viewing experience and maximizing the effectiveness of advertising.
[2234] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2235] Step 1:
[2236] Input: The user saves the video of the scene to be learned on the device.
[2237] How it works: The user uses a dedicated application to upload the video to be studied to the server.
[2238] Output: The learning video is saved on the server.
[2239] Step 2:
[2240] Input: The training video received by the server.
[2241] How it works: The server analyzes the video frame by frame and extracts features such as player movements, positioning, and ball trajectory, using generative AI models (Keras, TensorFlow).
[2242] Output: The extracted features are stored as tags in a database.
[2243] Step 3:
[2244] Input: The user saves game footage or advertising footage to the device.
[2245] How it works: Users use a dedicated application to upload game footage and advertising footage to the server.
[2246] Output: Match footage and advertising footage are saved on the server.
[2247] Step 4:
[2248] Input: Game footage and advertising footage received by the server, as well as pre-trained features.
[2249] How it works: The server analyzes the received video frame by frame and compares it with pre-trained features to pick out matching scenes. It then uses a generative AI model to confirm feature matches.
[2250] Output: Matching scenes are organized into a list and temporarily stored on the server.
[2251] Step 5:
[2252] Input: The user checks the extracted scene clips in a dedicated application.
[2253] How it works: While the user plays the clip, the device uses the camera and microphone to input the user's facial expressions, voice, and gestures into the emotion engine in real time.
[2254] Output: The emotion engine extracts the user's emotion data and sends it to the server.
[2255] Step 6:
[2256] Input: Emotion data received by the server, and a list of extracted scenes.
[2257] How it works: The server dynamically adjusts how the clip is presented based on the user's emotional data, for example adjusting the clip's playback speed if a positive emotion is detected, or revisiting a section of the clip if a negative emotion is detected.
[2258] Output: The adjusted scene is displayed to the user in real time.
[2259] Step 7:
[2260] Input: The user judges each clip as "true" or "false."
[2261] How it works: The user uses a dedicated application to send the results of their assessment of the clip to the server.
[2262] Output: The server records the user's judgment results in a database and updates the learning model to improve analysis accuracy in future.
[2263] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2264] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2265] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2266] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2267] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2268] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2269] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2270] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2271] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2272] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2273] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2274] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2275] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2276] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2277] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2278] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2279] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2280] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2281] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2282] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2283] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2284] The following is further disclosed regarding the above embodiment.
[2285] (Claim 1)
[2286] means for receiving a video of a scene to be learned and extracting features from the video;
[2287] a means for receiving a game video and extracting a scene that matches the feature amount from the game video;
[2288] a means for presenting the extracted scene to a user and allowing the user to determine whether the scene is a predetermined scene;
[2289] A means for recording the user's judgment result and updating the learning model to improve extraction accuracy from the next time onwards;
[2290] A system including:
[2291] (Claim 2)
[2292] 2. The system according to claim 1, further comprising means for extracting player movements, positioning, and ball trajectory as features of the scene to be learned.
[2293] (Claim 3)
[2294] 10. The system of claim 1, further comprising means for extracting features from videos taken at different distances and angles and extracting scenes based thereon.
[2295] "Example 1"
[2296] (Claim 1)
[2297] means for receiving a video of a scene to be learned and extracting features from the video;
[2298] a means for receiving a game video and extracting a scene that matches the feature amount from the game video;
[2299] a means for presenting the extracted scene to a user and allowing the user to determine whether the scene is a predetermined scene;
[2300] A means for recording the user's judgment result and updating the learning model to improve extraction accuracy from the next time onwards;
[2301] A method for analyzing the video frame by frame to extract player movements, positioning, and ball trajectory.
[2302] a means for storing the uploaded video in a storage system;
[2303] A means for selecting whether the judgment result is "correct" or "incorrect";
[2304] A system including:
[2305] (Claim 2)
[2306] 2. The system according to claim 1, further comprising means for extracting player movements, positioning, and ball trajectory as features of the scene to be learned.
[2307] (Claim 3)
[2308] 10. The system of claim 1, further comprising means for extracting features from videos taken at different distances and angles and extracting scenes based thereon.
[2309] "Application Example 1"
[2310] (Claim 1)
[2311] means for receiving a video of a work scene to be learned and extracting features from the video;
[2312] a means for receiving a manufacturing process video and extracting a scene that matches the feature amount from the manufacturing process video;
[2313] a means for presenting the extracted work scene to a user, and allowing the user to determine whether the scene is a predetermined work scene;
[2314] A means for recording the user's judgment results and updating the learning model to improve extraction accuracy from the next time onwards;
[2315] A method for analyzing situations where abnormalities or inefficiencies are suspected and organizing them into a list,
[2316] means for generating improvement suggestions based on the analysis results;
[2317] A system including:
[2318] (Claim 2)
[2319] 2. The system according to claim 1, further comprising means for extracting worker movements, tool use frequency, and work time as feature quantities of the work scene to be learned.
[2320] (Claim 3)
[2321] 2. The system according to claim 1, further comprising means for extracting features from manufacturing process videos taken at different distances and angles, and extracting work scenes based thereon.
[2322] "Example 2: Combining Emotion Engines"
[2323] (Claim 1)
[2324] means for receiving a video of a scene to be learned and extracting features from the video;
[2325] a means for receiving a game video and extracting a scene that matches the feature amount from the game video;
[2326] a means for presenting the extracted scene to a user and allowing the user to determine whether the scene is a predetermined scene;
[2327] A means for recording the user's judgment result and updating the learning model to improve extraction accuracy from the next time onwards;
[2328] means for recognizing a user's emotion and dynamically adjusting the content and presentation of the extracted scene based on the recognized emotion;
[2329] A system including:
[2330] (Claim 2)
[2331] 10. The system of claim 1, further comprising means for analyzing a user's facial expressions, voice, and gestures when a scene for emotion recognition is played back, and generating emotion data based thereon.
[2332] (Claim 3)
[2333] 2. The system according to claim 1, further comprising means for dynamically adjusting the playback speed of a scene and the display of detailed information according to the emotion based on the generated emotion data.
[2334] "Application example 2 when combining emotion engines"
[2335] (Claim 1)
[2336] means for receiving a video of a scene to be learned and extracting features from the video;
[2337] a means for receiving a game video and extracting a scene that matches the feature amount from the game video;
[2338] a means for presenting the extracted scene to a user and allowing the user to determine whether the scene is a predetermined scene;
[2339] A means for recording the user's judgment result and updating the learning model to improve extraction accuracy from the next time onwards;
[2340] means for recognizing a user's emotion while watching and dynamically adjusting the content and display method of the extracted scene based on the emotion;
[2341] A system including:
[2342] (Claim 2)
[2343] 2. The system according to claim 1, further comprising means for extracting player movements, positioning, and ball trajectory as features of the scene to be learned.
[2344] (Claim 3)
[2345] 10. The system of claim 1, further comprising means for extracting features from videos taken at different distances and angles and extracting scenes based thereon. [Explanation of symbols]
[2346] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving a video of a scene to be learned and extracting features from the video; a means for receiving a game video and extracting a scene that matches the feature amount from the game video; a means for presenting the extracted scene to a user and allowing the user to determine whether the scene is a predetermined scene; A means for recording the user's judgment result and updating the learning model to improve extraction accuracy from the next time onwards; A system including:
2. The system according to claim 1 , further comprising means for extracting player movements, positioning, and ball trajectory as features of the scene to be learned.
3. 2. The system according to claim 1, further comprising means for extracting features from videos taken at different distances and angles and extracting scenes based on the features.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A