system
The system addresses subjective scoring in artistic competitions by analyzing video and audio data with AI to provide fair and consistent artistic scores in real-time.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2026-03-10
AI Technical Summary
Conventional methods for evaluating artistic merit in competitions, such as figure skating and piano performances, rely heavily on subjective judgment, leading to inconsistent and unfair scoring.
A system that includes means for receiving, preprocessing, and analyzing video or audio data using AI models to calculate artistic scores, which are then displayed in real-time for fair and consistent evaluation.
Provides fair and consistent evaluation of artistic performances by quantitatively assessing movement and musical features, reducing variability and ensuring transparency in scoring.
Smart Images

Figure 2026041252000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, fair and objective scoring has become a necessity in competitions and performances that evaluate artistic merit, such as figure skating and piano competitions. However, conventional methods often rely on the subjective judgment and experience of judges, resulting in inconsistent scores. This can lead to unfair results for competitors and performers. The objective of this invention is to solve this problem and ensure consistency and fairness in the scoring. [Means for solving the problem]
[0005] The present invention relates to a system including a means for receiving video data or audio data, a means for pre-processing the received video data or audio data, a means for analyzing the pre-processed data and calculating an artistic score, and a means for displaying the calculated artistic score. Specifically, the system has the following configuration:
[0006] In the case of video data, video data including human movements is received, and the pre-processing means extracts the poses and movement features of the people and analyzes these features using an AI model to calculate an artistic score.In the case of audio data, audio data including musical performance is received, and the pre-processing means extracts musical features and analyzes these features using an AI model to calculate an artistic score.
[0007] The calculated artistic score is displayed in real time and provided to competitors and judges. In this way, the present invention allows for fair and consistent evaluation and reduces variability in artistic competitions and performances.
[0008] "Video data" refers to digital information in video format that includes people's actions and scenes.
[0009] "Audio Data" means digital information in audio form, including musical performances and sounds.
[0010] "Means for receiving" refers to a combination of hardware and software for acquiring digital data from the outside and incorporating it into the system.
[0011] "Pre-processing means" refers to a combination of hardware and software for converting received data into a form suitable for subsequent analysis steps.
[0012] "Means for analysis and calculation of artistic score" refers to algorithms and AI models that analyze pre-processed data and quantitatively evaluate the quality of performance or playing.
[0013] "Means for displaying" refers to a display device and associated software for visually outputting and providing the results of the analysis to a user or judge.
[0014] "Extracting a person's poses and movement characteristics" refers to the process of extracting characteristics such as the joint positions of the human body, movement patterns, jumps, and spins from video data.
[0015] "Extracting musical features" refers to the process of extracting musical elements such as pitch, intensity, tempo, and rhythm from audio data.
[0016] "AI model" refers to an algorithm based on machine learning and deep learning technologies used for data analysis.
[0017] "Real-time" means that results are displayed with only a short delay after the data is received. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The system of the present invention analyzes video or audio data in real time to provide fair and consistent evaluation in artistic scoring competitions. The system has a series of functions for receiving, preprocessing, and analyzing video or audio data, and displaying the resulting artistic scores. The specific operation of the system is described below.
[0040] System Overview
[0041] This system consists of a server, a terminal, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user.
[0042] 1. Data collection
[0043] Users host events such as figure skating and piano competitions.
[0044] A server receives video data (e.g., video of a figure skating performance) and audio data (e.g., audio of a piano performance) from an event in real time.
[0045] System operation details
[0046] 1. Data Collection
[0047] A server receives a live stream of a figure skating or piano competition.
[0048] In the case of video data, the server acquires data in real time from cameras and other video input sources.
[0049] For audio data, the server captures data in real time from a microphone or other audio input source.
[0050] 2. Data Preprocessing
[0051] The server converts the received video and audio data into an easy-to-understand format.
[0052] In the case of figure skating, motion capture and image recognition technology are used to extract a person's poses and movement characteristics.
[0053] In the case of piano performance, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0054] 3. Data Analysis
[0055] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[0056] In figure skating, the height of the jump, the number of rotations, and the speed of the spins are evaluated.
[0057] Piano performance is assessed on accuracy of tone, expressiveness, and rhythmic consistency.
[0058] The analyzed results (art scores) are stored in a database.
[0059] 4. Providing results
[0060] The terminal displays the analysis results (art scores) in real time.
[0061] The terminal is equipped with a display device for displaying information to competitors and judges.
[0062] The server records the results and makes them available for later analysis and review.
[0063] Specific examples
[0064] figure skating
[0065] 1. A user hosts a figure skating competition.
[0066] 2. The server receives live footage of the competition and collects data on jumps and spins.
[0067] 3. The server preprocesses the collected video data to extract information such as jump height and spin rotations.
[0068] 4. The server analyzes the preprocessed data using an AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[0069] 5. The device displays the analysis results in real time and provides them to the competitors and judges.
[0070] Piano performance
[0071] 1. A user hosts a piano competition.
[0072] 2. The server receives the audio of the competition performance.
[0073] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[0074] 4. The server analyzes the preprocessed data using an AI model to evaluate the sound characteristics and technical accuracy of the performance.
[0075] 5. The device displays the analysis results in real time and provides them to the user and judges.
[0076] In this way, the system of the present invention is able to provide fair and consistent grading for artistic competitions and reduce variability in grading.
[0077] The processing flow will be explained below.
[0078] Processing flow and each processing step
[0079] Step 1: Receiving Data
[0080] A user hosts a figure skating competition or a piano competition.
[0081] A server receives video or audio data from an event in real time.
[0082] In the case of video data, the data is obtained from a camera or video input source.
[0083] For audio data, data is obtained from a microphone or audio input source.
[0084] Step 2: Save your data
[0085] Create a directory for the server to temporarily store the data it receives.
[0086] The server temporarily saves the received data stream in the "raw_data" directory.
[0087] Step 3: Preprocessing the data
[0088] The server reads the saved data from the "raw_data" directory.
[0089] In the case of video data, motion capture and image recognition technology are used to extract people's poses and movement characteristics.
[0090] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0091] The server saves the preprocessed data in the "processed_data" directory.
[0092] Step 4: Analyze the data
[0093] The server reads the preprocessed data from the "processed_data" directory.
[0094] The server inputs the preprocessed data into an AI model to calculate the artistic score.
[0095] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[0096] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[0097] The server stores the analysis results (art scores) in a database.
[0098] Step 5: Delivering results
[0099] The terminal obtains the analysis results (art scores) in real time.
[0100] The terminal provides a display device for displaying the analysis results to the competitors and judges.
[0101] The user can check the displayed analysis results and use them to evaluate the competition or performance.
[0102] The server records the analysis results and stores them in a database for future analysis and review.
[0103] In this way, with each step working together, the system of the present invention provides fair and consistent evaluation of artistic competitions and performances.
[0104] Example 1
[0105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0106] Conventional evaluation methods for artistic competitions often rely on the subjective judgment and experience of judges, resulting in a lack of fairness and consistency. Real-time evaluation is also difficult, resulting in variations and delays in evaluation. The present invention aims to solve these problems and provide a system that provides fair and consistent evaluation in real time.
[0107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0108] In this invention, the server includes a means for receiving video data or audio data, a means for preprocessing the received video data or audio data, and a means for inputting the preprocessed data into a generative AI model to calculate artistic scores, thereby enabling the server to analyze the video data or audio data in real time and provide fair and consistent evaluations.
[0109] "Video Data" refers to visual information of an action or scene captured from a camera or other video input device.
[0110] "Audio Data" refers to sound information collected from a microphone or other audio input device.
[0111] "Means for receiving" refers to the technical means by which the server retrieves video and audio data from a network or device.
[0112] "Pre-processing means" refers to technical means for converting received video and audio data into an understandable format. Examples include motion capture and audio signal processing techniques.
[0113] A "generative AI model" is a model that uses artificial intelligence technology to analyze data and generate scores or ratings for specific purposes.
[0114] "Means for calculating artistic scores" refers to the technical means for inputting preprocessed data into a generative AI model and conducting a quantitative evaluation.
[0115] "Means for displaying" refers to the technical means for visually presenting the artistic score, which is the result of the analysis, to users and judges.
[0116] "Personal action" refers to a figure skating routine or other artistic movement performed by a particular person.
[0117] "Extracting pose and movement features" refers to the act of analyzing specific postures and movements from video data and extracting related features.
[0118] "Musical performance" refers to the act of performing a musical piece using instruments and voices.
[0119] "Extracting musical features" refers to the act of analyzing attributes such as pitch, intensity, and tempo of sound from audio data and extracting features.
[0120] The system of the present invention uses video and audio data analysis technology to provide fair and consistent artistic evaluation in real time. The system is mainly composed of three elements: a server, a terminal, and a user.
[0121] server
[0122] Data reception: The server first receives video and audio data from events such as figure skating and piano competitions. Specifically, it acquires streaming data via cameras and microphones using RTSP (Real-Time Streaming Protocol) or similar.
[0123] Data preprocessing: The server then preprocesses the received data. For video data, motion capture and image recognition technologies are used to extract human poses and movement characteristics. This is done using libraries such as OpenPose. For audio data, FFT (Fast Fourier Transform) is used to extract musical features such as pitch, intensity, and tempo.
[0124] Data Analysis: The pre-processed data is then fed into a generative AI model. The server uses deep learning models such as TENSORFLOW® or PyTorch to analyze this data and calculate artistic scores. In the case of figure skating, scores are generated based on factors such as jump height, number of rotations, and spin speed. In the case of piano playing, scores are evaluated based on pitch accuracy, tempo consistency, and expressiveness.
[0125] Result storage: The analysis results (art scores) are stored in a database by the server, allowing for later analysis and review. This can be done using a SQL or NoSQL database.
[0126] Terminal
[0127] Display of results: The device displays the analysis results obtained from the server in real time. An interface is provided to visually present the results to competitors and judges. Specifically, the results are displayed using a web browser or a mobile app.
[0128] Accepting user input: The terminal also provides an interface for accepting user input, allowing operations such as starting and stopping a competition and checking results.
[0129] User
[0130] Hosting an event: A user hosts an event such as a figure skating or piano competition. They distribute entry forms to recruit participants, schedule the competition, and notify them.
[0131] Specific examples
[0132] In the case of figure skating
[0133] 1. A user hosts a figure skating competition.
[0134] Example: A user posts a competition and gathers participants online.
[0135] 2. The server receives live footage of the competition and collects data on jumps and spins.
[0136] Specific operation: The server receives video streaming data from the camera and extracts motion features using OpenPose.
[0137] 3. The server analyzes the preprocessed data using a generative AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[0138] Example prompt: "Analyze video of a figure skating competition and measure the jump height, number of rotations, and spin speed. Calculate an artistic score based on this data."
[0139] 4. The device displays the analysis results in real time and provides them to the competitors and judges.
[0140] Example: Results are displayed instantly on large screens at the stadium.
[0141] For piano performances
[0142] 1. A user hosts a piano competition.
[0143] Specific actions: Recruit participants using an entry form.
[0144] 2. The server receives the audio of the competition performance.
[0145] Example: A server captures audio data from a microphone in real time.
[0146] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[0147] Example: Extracting sound characteristics using FFT.
[0148] 4. The server analyzes the preprocessed data using a generative AI model to evaluate the sound characteristics and technical accuracy of the performance.
[0149] 5. The device displays the analysis results in real time and provides them to the user and judges.
[0150] Example: Analysis results are displayed instantly on smartphone apps and web apps.
[0151] In this way, the system of the present invention is able to provide fair and consistent grading for artistic competitions and reduce variability in grading.
[0152] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0153] Step 1: Collect data
[0154] A user hosts an event such as a figure skating or piano competition. Specifically, the user prepares an entry form, recruits participants, and sets and notifies the schedule of the competition. This completes the preparations for the competition.
[0155] The server receives video and audio data from the event in real time. Specific inputs include streaming video and audio data captured using a camera or microphone. The raw data is then stored on the server.
[0156] Step 2: Preprocessing the data
[0157] The server preprocesses the video and audio data it receives. Specifically, the server receives raw data and converts it into a format that is easy to analyze.
[0158] For video data, the server uses motion capture and image recognition technology to extract poses and movement characteristics of people, and specific outputs, such as jump height and spin rotations, are extracted using libraries such as OpenPose.
[0159] For audio data, the server uses FFT (Fast Fourier Transform) to extract musical features such as pitch, intensity, and tempo. Specific outputs include pitch accuracy, intensity distribution, and tempo fluctuation data.
[0160] Step 3: Analyze the data
[0161] The server inputs the preprocessed data into a generative AI model to calculate the artistic score. Specific inputs include preprocessed motion data and musical features. This is then input into an AI model such as TensorFlow or PyTorch.
[0162] In the case of figure skating, the server evaluates the height of jumps, number of rotations, speed of spins, etc. The output is a score for each element and an overall artistic score.
[0163] For piano performances, the server evaluates pitch accuracy, tempo consistency, and expressiveness, producing an output that measures the overall performance's technical accuracy and artistic evaluation score.
[0164] Step 4: Save the results
[0165] The server stores the analysis results (art scores) in a database. Specific inputs include the generated art scores and scores for each evaluation item. These are recorded in an SQL or NoSQL database.
[0166] As an output, the analysis results are stored in a format that can be used for later analysis and review, for example to analyze historical and trending results.
[0167] Step 5: View the results
[0168] The device displays the analysis results (art scores) in real time. Specific inputs include analysis results obtained from the server, which are then displayed on the user interface of a web browser or mobile app.
[0169] The output allows athletes and judges to check the evaluation results in real time. For example, the results may be displayed on a large screen at the competition venue or instantly on a smartphone app.
[0170] Step 6: User feedback
[0171] The user checks the evaluation results through the terminal and provides feedback. Specific inputs include feedback data from the user, which is sent to the server and stored.
[0172] As an output, the feedback is used to refine the system and adjust the evaluation criteria for the next iteration, thereby improving the overall accuracy of the system and the user experience.
[0173] (Application example 1)
[0174] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0175] With conventional methods, evaluation of live performances in physical venues is often subjective, making it difficult to achieve fair and consistent evaluations. Furthermore, the delay in real-time evaluation feedback often makes performers and audiences feel that the evaluations lack transparency and immediacy.
[0176] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0177] In this invention, the server includes means for receiving video data or audio data, means for pre-processing the received video data or audio data, means for analyzing the pre-processed data and calculating an artistic score, means for displaying the calculated artistic score, and means for capturing video and audio data of live performances through cameras and microphones in the store, thereby enabling fair and consistent real-time evaluation of live performances in physical stores.
[0178] "Video data" refers to dynamic or static image information acquired through a visual device such as a camera.
[0179] "Audio data" refers to an audio signal obtained through an audio device such as a microphone.
[0180] "Preprocessing" refers to a series of data manipulation steps that convert the received raw data into a format that is easier to analyze.
[0181] "Art score calculation" refers to the process of analyzing pre-processed data and generating a score based on specific evaluation criteria.
[0182] "Display means" refers to a display device or software interface that visually presents the analysis results to the user.
[0183] "In-store cameras" refer to video capture devices installed in physical stores to capture video data of live performances.
[0184] "In-store microphones" refer to audio capture devices installed in physical stores to capture audio data from live performances.
[0185] The system embodying this invention evaluates live performance in a physical store in real time and provides fair and consistent scoring. The system mainly uses a server, in-store cameras and microphones, a smartphone app, and a display.
[0186] Hardware and Software Configuration
[0187] 1. Server:
[0188] Role: Receiving, preprocessing, analyzing and storing data.
[0189] Software used:
[0190] Motion capture technology: OpenPose, etc.
[0191] Audio signal processing technology: librosa, etc.
[0192] AI analysis model: TensorFlow or PyTorch
[0193] 2. In-store cameras:
[0194] Role: Capture video data of the performance.
[0195] Example: High-definition video camera
[0196] 3. In-store microphones:
[0197] Role: Captures audio data of the performance.
[0198] Example: High-sensitivity microphone
[0199] 4. Smartphone App:
[0200] Role: Displays the score and accepts user input.
[0201] Example: An app for ANDROID (registered trademark) or iOS
[0202] 5. Display:
[0203] Role: Displaying evaluation results in-store.
[0204] Example: Large displays in stores
[0205] A natural language description of the program's operation
[0206] The server captures video data of the performance in real time through in-store cameras, and simultaneously captures audio data through in-store microphones. Motion capture technology is used to extract the poses and movement characteristics of the performers from the received video data, and audio signal processing technology is used to extract musical features such as pitch, intensity, and tempo from the audio data. The extracted features are input into an AI analysis model to calculate the performance's artistic score. The calculated artistic score is saved on the server and displayed in real time on a smartphone app and on in-store displays.
[0207] Specific examples
[0208] For example, imagine a live music concert is being held in a store. When the live performance begins, cameras in the store collect video data and microphones collect audio data, which are then sent to a server. The server preprocesses the received video data and analyzes the performers' movement characteristics using OpenPose. Musical features are extracted from the audio data using librosa, and these are input into an AI model built with TensorFlow and PyTorch. The model generates a score based on evaluation criteria, and the score is displayed in real time on a smartphone app or display.
[0209] Prompt Sentence Examples
[0210] Please use an AI model to analyze the beauty of jumps and spins in figure skating performance footage and provide a real-time evaluation score. Please enter the video data below. Please display the analysis results.
[0211] <Video data>
[0212] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0213] Step 1:
[0214] Data collection
[0215] When a live event begins, cameras and microphones in the store capture video and audio data of the performance in real time. These data are sent to a server. The cameras provide video data (e.g., the performers' movements) and the microphones provide audio data (e.g., the musical performance).
[0216] Input: Live performance video and audio data
[0217] Output: Video and audio data sent to the server
[0218] Step 2:
[0219] Data Preprocessing
[0220] The server preprocesses the received video data using OpenPose to extract human poses and movement features, and preprocesses the audio data using librosa to extract musical features such as pitch, intensity, and tempo, resulting in an analyzable dataset.
[0221] Input: Video and audio data sent to the server
[0222] Output: Preprocessed motion feature data and musical feature data
[0223] Step 3:
[0224] Data analysis
[0225] The server inputs the preprocessed data into an AI analysis model (e.g., using TensorFlow or PyTorch) to calculate the artistic score of the performance. The analysis model uses features such as the accuracy of the movements, the rhythm of the music, and the tempo as evaluation criteria.
[0226] Input: Preprocessed motion feature data and musical feature data
[0227] Output: Calculated art score
[0228] Step 4:
[0229] Saving the results
[0230] The server stores the calculated artistic scores in a database, which can be used for later review and analysis.
[0231] Input: Calculated art score
[0232] Output: Art points stored in the database
[0233] Step 5:
[0234] Displaying the results
[0235] The device displays the analysis results in real time, i.e., artistic scores, on a smartphone app and on displays in the store, allowing performers and audience members to immediately check the evaluation results.
[0236] Input: Art points stored in the database
[0237] Output: Smartphone app and artwork displayed on the screen
[0238] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0239] The system of the present invention includes a series of processes from receiving video or audio data to pre-processing, analysis, and display, and is combined with an emotion engine that recognizes the user's emotions. This system not only enables fair and consistent evaluation in artistic competitions and performances, but also takes into account the user's emotional responses.
[0240] System Overview
[0241] This system consists of a server, a terminal, an emotion engine, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user. The emotion engine recognizes the user's emotions and fine-tunes the evaluation based on their emotional response.
[0242] 1. Data collection
[0243] Users host figure skating and piano competitions.
[0244] The server receives video data (e.g., video of a figure skating performance) and audio data (e.g., audio of a piano performance) from the event in real time.
[0245] System operation details
[0246] 1. Data Collection
[0247] The server receives the video and audio data in real time.
[0248] In the case of video data, the data is obtained from a camera or video input source.
[0249] For audio data, data is obtained from a microphone or audio input source.
[0250] 2. Data Preprocessing
[0251] The server preprocesses the received video and audio data.
[0252] In the case of video data, motion capture and image recognition technology are used to extract a person's pose and movement characteristics.
[0253] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0254] 3. Data Analysis
[0255] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[0256] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[0257] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[0258] The analysis results (art scores) obtained are stored in a database.
[0259] 4. Emotion recognition
[0260] An emotion engine analyzes the facial expressions, vocal or body language of athletes, performers and spectators to identify their emotional state.
[0261] The server uses the output of the emotion engine to fine-tune the artistic score.
[0262] 5. Providing results
[0263] The terminal obtains the analysis results (art scores) in real time.
[0264] The terminal displays the analysis results to the competitor and judges.
[0265] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[0266] The server records the analysis results and stores them in a database for future analysis and review.
[0267] Specific examples
[0268] figure skating
[0269] 1. A user hosts a figure skating competition.
[0270] 2. The server receives live footage of the competition and collects data on jumps and spins.
[0271] 3. The server preprocesses the collected video data to extract information such as jump height and spin rotations.
[0272] 4. The server analyzes the preprocessed data using an AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[0273] 5. The emotion engine analyzes the facial expressions of the audience and judges and measures the level of emotion based on each element.
[0274] 6. The server fine-tunes the artistic score based on the output of the emotion engine.
[0275] 7. The device displays the analysis results in real time and provides them to the competitors and judges.
[0276] Piano performance
[0277] 1. A user hosts a piano competition.
[0278] 2. The server receives the audio of the competition performance.
[0279] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[0280] 4. The server analyzes the preprocessed data using an AI model to evaluate the sound characteristics and technical accuracy of the performance.
[0281] 5. The emotion engine analyzes the facial expressions of the audience and judges and measures the level of emotion based on each element.
[0282] 6. The server fine-tunes the artistic score based on the output of the emotion engine.
[0283] 7. The device displays the analysis results in real time and provides them to the user and judges.
[0284] In this way, each step works together to enable the system of the present invention to provide fair and consistent evaluations, and also to take into account the user's emotional responses.
[0285] The processing flow will be explained below.
[0286] Processing flow and each processing step
[0287] Step 1: Receiving Data
[0288] A user hosts a figure skating competition or a piano competition.
[0289] A server receives video or audio data from an event in real time.
[0290] In the case of video data, the data is obtained from a camera or video input source.
[0291] For audio data, data is obtained from a microphone or audio input source.
[0292] Step 2: Save your data
[0293] Create a directory for the server to temporarily store the data it receives.
[0294] The server temporarily saves the received data stream in the "raw_data" directory.
[0295] Step 3: Preprocessing the data
[0296] The server reads the saved data from the "raw_data" directory and performs preprocessing.
[0297] In the case of video data, motion capture and image recognition technology are used to extract people's poses and movement characteristics.
[0298] Motion capture is a technology that identifies the joint positions of a person in a video.
[0299] Image recognition technology is a technology that recognizes specific movements (e.g., jumps and spins).
[0300] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0301] Pitch is a feature that indicates the pitch of a sound.
[0302] Intensity is a feature that indicates the loudness of a sound.
[0303] Tempo is a characteristic that indicates the speed of a performance.
[0304] The server saves the preprocessed data in the "processed_data" directory.
[0305] Step 4: Analyze the data
[0306] The server reads the preprocessed data from the "processed_data" directory.
[0307] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[0308] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[0309] The jump height refers to the height from the start point of the jump to the landing point.
[0310] The number of rotations refers to the number of rotations during the jump.
[0311] The rotation speed of the spin indicates the speed of rotation during the spin.
[0312] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[0313] The accuracy of the sound is evaluated by whether the sound is produced according to the musical score.
[0314] Expressiveness is measured by the performer's richness of emotion and expression.
[0315] Rhythmic consistency is measured by whether the performance maintains a consistent tempo.
[0316] The server stores the analysis results (art scores) in a database.
[0317] Step 5: Emotion Recognition
[0318] The emotion engine analyzes the facial expressions, voice or body language of athletes and spectators to identify their emotional state.
[0319] Facial expression analysis is a technology that analyzes facial expressions from camera footage and identifies emotions.
[0320] Voice analysis is a technology that analyzes the tone and strength of a voice to identify emotions.
[0321] Body language analysis is a technique that analyzes gestures and postures to identify emotions.
[0322] The server uses the output of the emotion engine to fine-tune the artistic score.
[0323] The emotional score is a numerical representation of the user's emotional response.
[0324] To fine-tune the artistic score, the emotional score is taken into account in the final evaluation.
[0325] Step 6: Delivering results
[0326] The terminal obtains the analysis results (art scores) in real time.
[0327] The terminal provides a display device for displaying the analysis results to the competitors and judges.
[0328] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[0329] The server records the analysis results and stores them in a database for future analysis and review.
[0330] In this way, each step works together to enable the system of the present invention to provide fair and consistent evaluations, and also to take into account the user's emotional responses.
[0331] Example 2
[0332] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0333] Conventional evaluation systems for artistic competitions and performances have difficulty providing fair and consistent evaluations. In particular, they lack systems that reflect the emotional reactions of audiences and judges in their evaluations, resulting in a lack of objectivity in the evaluations. Furthermore, they lack the ability to provide evaluations in real time and to store evaluation results.
[0334] The identification processing by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving video data or audio data, means for preprocessing the received video data or audio data, means for analyzing the preprocessed data and calculating an artistic score, means for recognizing the emotions of the contestants and spectators, means for fine-tuning the artistic score based on the emotion recognition results, and means for displaying the calculated artistic score. This makes it possible to provide fair and consistent evaluations and realize evaluations that take into account the emotional reactions of spectators and judges.
[0335] "Video data" refers to data that includes image information obtained from a camera, video input source, or the like.
[0336] "Audio data" refers to data containing sound information obtained from a microphone, audio input source, or the like.
[0337] "Preprocessing" refers to performing initial processing to extract specific features from received video and audio data.
[0338] "Analysis" refers to the evaluation of pre-processed data using AI models or other means to produce a specific result (e.g., artistic score).
[0339] "Artistic Score" is a score calculated based on the received video and audio data, or the emotional response of the competitors and spectators.
[0340] "Emotion recognition" is the process of analyzing the facial expressions, voice, body language, etc. of athletes and spectators to identify their emotional state.
[0341] "Fine-tuning based on emotion recognition results" means adjusting the aforementioned analysis results (artistic score) based on the data obtained through emotion recognition.
[0342] "Display" refers to the visual presentation of the analyzed results and fine-tuned artistic points.
[0343] The system of the present invention realizes a series of processes from receiving video or audio data, to preprocessing, analysis, display, and further fine-tuning evaluation by emotion recognition. This section explains the specific implementation method.
[0344] This system mainly consists of a server, a terminal, an emotion engine, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user. The emotion engine recognizes the user's emotions and fine-tunes the evaluation based on their emotional response.
[0345] Data collection
[0346] First, a user hosts an artistic competition or performance event (e.g., figure skating or piano competition). The server receives real-time video and audio data from the event. Video data comes from cameras and other video input sources, and audio data comes from microphones and other audio input sources.
[0347] Data Preprocessing
[0348] The server preprocesses the received video and audio data. For video data, motion capture and image recognition technology are used to extract the poses and movement characteristics of the person. For audio data, audio signal processing technology is used to extract musical features such as the pitch, intensity, and tempo of the sound. For example, the height of a figure skater's jump or the number of rotations in a spin, or the accuracy and rhythmic consistency of a piano performance can be obtained.
[0349] Data analysis
[0350] The server inputs the preprocessed data into an AI model to calculate artistic scores. In the case of video data, the AI model evaluates jump height, number of rotations, spin speed, etc., and in the case of audio data, it evaluates sound accuracy, expressiveness, rhythmic consistency, etc. The analysis results (artistic scores) obtained are stored in a database.
[0351] emotion recognition
[0352] The emotion engine analyzes the facial expressions, voice, and body language of competitors, performers, and audience members to identify their emotional state. The server uses the output of the emotion engine to fine-tune the artistic score, enabling a fair evaluation that reflects not only pure technical evaluation but also the emotional reactions of the audience and judges.
[0353] Providing results
[0354] The device acquires the analysis results (artistic scores) in real time and displays them to the competitors and judges. The user can check the displayed analysis results and use them to evaluate the competition and performance. The server also records the analysis results and stores them in a database for future analysis and review.
[0355] Specific examples
[0356] In the case of figure skating: A user hosts a competition, and the server receives live footage of the competition and collects data on jumps and spins. The server preprocesses this data, analyzes the jump height and spin rate using an AI model, and calculates and adjusts the artistic score. The results are displayed on the device in real time and provided to competitors and judges.
[0357] In the case of piano performance: A user hosts a piano competition, and the server receives the audio of the competition performance. The server preprocesses the audio data and analyzes pitch, intensity, tempo, etc. using an AI model to evaluate the accuracy of the sound and the technical accuracy of the performance. The emotion engine analyzes the facial expressions of the audience and judges and fine-tunes the artistic score based on their emotions. The results are displayed on the device in real time and provided to the user and judges.
[0358] Example prompt sentence:
[0359] Explain the processing steps of a system that evaluates the height of figure skating jumps and the number of rotations in spins, and adjusts the final artistic score taking into account the emotional response of the audience.
[0360]
[0361] The system analyzes musical characteristics such as pitch, intensity, and tempo during piano performances, and fine-tunes its evaluation based on the emotional responses of the audience and judges.
[0362] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0363] Step 1: Collect data
[0364] A user hosts events such as figure skating and piano competitions.
[0365] The server receives video or audio data from the event in real time. Specifically, video data is obtained from a camera or video input, and audio data is obtained from a microphone or audio input.
[0366] Input: Video data from a camera or video input, audio data from a microphone or audio input
[0367] Output: Raw data stored on the server
[0368] Step 2: Data Preprocessing
[0369] The server preprocesses the video and audio data received.
[0370] In the case of video data, motion capture and image recognition technology are used to extract a person's pose and movement characteristics.
[0371] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0372] Input: Received raw video and audio data
[0373] Output: Preprocessed feature data (e.g., person pose data, sound characteristic data)
[0374] Step 3: Data analysis
[0375] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[0376] When using video data, the AI model evaluates factors such as jump height, number of rotations, and spin speed.
[0377] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[0378] Input: Preprocessed feature data
[0379] Output: Art score calculated by the AI model
[0380] Step 4: Emotion Recognition
[0381] The emotion engine analyzes the facial expressions, voice or body language of athletes, performers and spectators to identify their emotional state.
[0382] The server fine-tunes the artistic score based on the output of the emotion engine.
[0383] Input: Emotional data obtained from the camera or microphone (facial expressions, voice, body language)
[0384] Output: Emotional state analyzed by emotion engine, fine-tuned artistic score
[0385] Step 5: Delivering results
[0386] The terminal obtains the analysis results (art scores) in real time.
[0387] The terminal displays the analysis results to the competitors and judges.
[0388] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[0389] The server records the analysis results and stores them in a database for future analysis and review.
[0390] Input: Fine-tuned art scores, emotion recognition results
[0391] Output: Final evaluation results displayed in real time on the device, evaluation results stored in the database
[0392] (Application example 2)
[0393] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0394] Currently, there are almost no systems in physical stores that can grasp customer emotions and behavior in real time and provide appropriate responses based on that information. This makes it difficult to provide services that immediately reflect customer needs and emotions, limiting the improvement of customer satisfaction. It is also difficult to introduce a fair and consistent evaluation system.
[0395] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving video data or audio data, means for preprocessing the received video data or audio data, means for analyzing the preprocessed data and calculating an artistic score, means for displaying the calculated artistic score, and means for recognizing the user's emotions and fine-tuning the evaluation based on the emotional response. This makes it possible to analyze customer emotions and behavior in real time in physical stores and provide appropriate feedback to store staff immediately. Furthermore, using an emotion recognition engine enables a fair and consistent evaluation system and improves customer satisfaction.
[0396] "Video data" refers to image or video data acquired through a camera, video device, or the like.
[0397] "Audio data" refers to sound or voice data acquired through a microphone, recording device, or the like.
[0398] "Receiving means" refers to a function or system that allows a server or terminal to receive video data or audio data from the outside.
[0399] "Preprocessing means" refers to a function or system that processes received video data and audio data and converts them into a format suitable for analysis.
[0400] "Analysis means" refers to a function or system for calculating evaluation points based on preprocessed data.
[0401] "Artistic score" is a numerical evaluation of video and audio, and is an index used to evaluate technical and artistic aspects.
[0402] "Display means" refers to a function or system for visually displaying the calculated evaluation scores and analysis results.
[0403] "Emotion recognition" is a technology that identifies a user's emotional state based on data such as facial expressions and voice.
[0404] "Fine-tuning of the evaluation" is the process of adjusting the evaluation score calculated based on the emotion recognition results.
[0405] A "camera" is a photographing device for acquiring video data.
[0406] A "microphone" is a recording device for capturing audio data.
[0407] An "emotion engine" is software or hardware for analyzing a user's emotional state.
[0408] A "physical store" is a sales location set up in a physical location where customers visit in person to make purchases.
[0409] A "store clerk" is an employee who provides services to customers in a physical store.
[0410] The system of the present invention recognizes customer emotions and provides feedback to store clerks in real time. This system consists of a server, a camera, a microphone, a terminal, and an emotion engine. Specific hardware and software used include OpenCV, Keras, librosa, etc.
[0411] The system is configured as follows:
[0412] 1. Data Collection:
[0413] The server receives video and audio data from cameras and microphones installed in the store. The cameras capture customers' facial expressions in real time, and the microphones record their voices.
[0414] 2. Data preprocessing:
[0415] The received video and audio data are preprocessed by the server. The video data undergoes face recognition and facial expression analysis using OpenCV. The audio data undergoes audio feature extraction (e.g., MFCC) using librosa.
[0416] 3. Data Analysis:
[0417] The pre-processed data is fed into a generative AI model using Keras to analyze the customer's emotional state. Emotions are recognized from facial and vocal characteristics, and specific emotional responses (e.g., joy, anger, sadness, etc.) are identified by an emotion engine.
[0418] 4. Providing results:
[0419] The server displays the analysis results on the terminal in real time. The store clerk can then use this information to respond appropriately to the customer. For example, if the customer is angry, a more careful response appropriate to the situation will be recommended. If the customer is satisfied, the usual service will be provided.
[0420] As a concrete example, we will show an example of inputting a prompt sentence into a generative AI model.
[0421] Prompt: "Implement a system in a brick-and-mortar store that recognizes customer emotions and instructs store associates on appropriate responses based on facial and voice data. Key scenarios to consider include when a customer is angry or sad."
[0422] The above configuration and procedures make it possible to improve the customer experience in physical stores, and are expected to increase customer satisfaction.
[0423] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0424] Step 1: Collect data
[0425] The server receives video and audio data of customers in the store via a camera and microphone. Specifically, the camera captures the customer's face and body movements, and the microphone records the customer's voice. The input is video data from the camera and audio data from the microphone, which are then sent to the server.
[0426] Step 2: Preprocessing the data
[0427] The server preprocesses the received video and audio data. Specifically, it uses OpenCV to recognize faces and facial expressions from the video data and normalizes the images. For audio data, it uses librosa to convert the audio signal into features (e.g., MFCCs). The input is raw video and audio data, and the output is preprocessed face image data and audio feature data.
[0428] Step 3: Analyze the data
[0429] The server analyzes the preprocessed data. Specifically, a generative AI model using Keras is used to estimate the emotional state from the facial image data and the emotional state of the voice from the voice feature data. The input is the preprocessed facial image data and voice feature data, and the output is the customer's emotional state (e.g., joy, anger, sadness, etc.).
[0430] Step 4: Applying the Emotion Engine
[0431] The server further processes the analysis results using an emotion engine to refine the accuracy. The emotion engine refines the overall emotion rating based on the analyzed emotional state. The input is the analyzed emotional state, and the output is a fine-tuned emotion rating.
[0432] Step 5: View the results
[0433] The server displays the adjusted emotion evaluation on the terminal in real time. Specifically, it notifies the store clerk of the analysis results and the fine-tuned emotion evaluation. The input is the fine-tuned emotion evaluation, and the output is a notification message displayed on the store clerk's terminal. For example, if the customer is angry, the message displayed is "The customer is angry. Please be careful how you respond."
[0434] Through these steps, the system can recognize customer emotions in real time in physical stores and provide immediate and appropriate feedback to store staff.
[0435] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0436] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0437] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0438] [Second embodiment]
[0439] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0440] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0441] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0442] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0443] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0444] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0445] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0446] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0447] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0448] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0449] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0450] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0451] The system of the present invention analyzes video or audio data in real time to provide fair and consistent evaluation in artistic scoring competitions. The system has a series of functions for receiving, preprocessing, and analyzing video or audio data, and displaying the resulting artistic scores. The specific operation of the system is described below.
[0452] System Overview
[0453] This system consists of a server, a terminal, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user.
[0454] 1. Data collection
[0455] Users host events such as figure skating and piano competitions.
[0456] A server receives video data (e.g., video of a figure skating performance) and audio data (e.g., audio of a piano performance) from an event in real time.
[0457] System operation details
[0458] 1. Data Collection
[0459] A server receives a live stream of a figure skating or piano competition.
[0460] In the case of video data, the server acquires data in real time from cameras and other video input sources.
[0461] For audio data, the server captures data in real time from a microphone or other audio input source.
[0462] 2. Data Preprocessing
[0463] The server converts the received video and audio data into an easy-to-understand format.
[0464] In the case of figure skating, motion capture and image recognition technology are used to extract a person's poses and movement characteristics.
[0465] In the case of piano performance, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0466] 3. Data Analysis
[0467] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[0468] In figure skating, the height of the jump, the number of rotations, and the speed of the spins are evaluated.
[0469] Piano performance is assessed on accuracy of tone, expressiveness, and rhythmic consistency.
[0470] The analyzed results (art scores) are stored in a database.
[0471] 4. Providing results
[0472] The terminal displays the analysis results (art scores) in real time.
[0473] The terminal is equipped with a display device for displaying information to competitors and judges.
[0474] The server records the results and makes them available for later analysis and review.
[0475] Specific examples
[0476] figure skating
[0477] 1. A user hosts a figure skating competition.
[0478] 2. The server receives live footage of the competition and collects data on jumps and spins.
[0479] 3. The server preprocesses the collected video data to extract information such as jump height and spin rotations.
[0480] 4. The server analyzes the preprocessed data using an AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[0481] 5. The device displays the analysis results in real time and provides them to the competitors and judges.
[0482] Piano performance
[0483] 1. A user hosts a piano competition.
[0484] 2. The server receives the audio of the competition performance.
[0485] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[0486] 4. The server analyzes the preprocessed data using an AI model to evaluate the sound characteristics and technical accuracy of the performance.
[0487] 5. The device displays the analysis results in real time and provides them to the user and judges.
[0488] In this way, the system of the present invention is able to provide fair and consistent grading for artistic competitions and reduce variability in grading.
[0489] The processing flow will be explained below.
[0490] Processing flow and each processing step
[0491] Step 1: Receiving Data
[0492] A user hosts a figure skating competition or a piano competition.
[0493] A server receives video or audio data from an event in real time.
[0494] In the case of video data, the data is obtained from a camera or video input source.
[0495] For audio data, data is obtained from a microphone or audio input source.
[0496] Step 2: Save your data
[0497] Create a directory for the server to temporarily store the data it receives.
[0498] The server temporarily saves the received data stream in the "raw_data" directory.
[0499] Step 3: Preprocessing the data
[0500] The server reads the saved data from the "raw_data" directory.
[0501] In the case of video data, motion capture and image recognition technology are used to extract people's poses and movement characteristics.
[0502] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0503] The server saves the preprocessed data in the "processed_data" directory.
[0504] Step 4: Analyze the data
[0505] The server reads the preprocessed data from the "processed_data" directory.
[0506] The server inputs the preprocessed data into an AI model to calculate the artistic score.
[0507] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[0508] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[0509] The server stores the analysis results (art scores) in a database.
[0510] Step 5: Delivering results
[0511] The terminal obtains the analysis results (art scores) in real time.
[0512] The terminal provides a display device for displaying the analysis results to the competitors and judges.
[0513] The user can check the displayed analysis results and use them to evaluate the competition or performance.
[0514] The server records the analysis results and stores them in a database for future analysis and review.
[0515] In this way, with each step working together, the system of the present invention provides fair and consistent evaluation of artistic competitions and performances.
[0516] Example 1
[0517] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0518] Conventional evaluation methods for artistic competitions often rely on the subjective judgment and experience of judges, resulting in a lack of fairness and consistency. Real-time evaluation is also difficult, resulting in variations and delays in evaluation. The present invention aims to solve these problems and provide a system that provides fair and consistent evaluation in real time.
[0519] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0520] In this invention, the server includes a means for receiving video data or audio data, a means for preprocessing the received video data or audio data, and a means for inputting the preprocessed data into a generative AI model to calculate artistic scores, thereby enabling the server to analyze the video data or audio data in real time and provide fair and consistent evaluations.
[0521] "Video Data" refers to visual information of an action or scene captured from a camera or other video input device.
[0522] "Audio Data" refers to sound information collected from a microphone or other audio input device.
[0523] "Means for receiving" refers to the technical means by which the server retrieves video and audio data from a network or device.
[0524] "Pre-processing means" refers to technical means for converting received video and audio data into an understandable format. Examples include motion capture and audio signal processing techniques.
[0525] A "generative AI model" is a model that uses artificial intelligence technology to analyze data and generate scores or ratings for specific purposes.
[0526] "Means for calculating artistic scores" refers to the technical means for inputting preprocessed data into a generative AI model and conducting a quantitative evaluation.
[0527] "Means for displaying" refers to the technical means for visually presenting the artistic score, which is the result of the analysis, to users and judges.
[0528] "Personal action" refers to a figure skating routine or other artistic movement performed by a particular person.
[0529] "Extracting pose and movement features" refers to the act of analyzing specific postures and movements from video data and extracting related features.
[0530] "Musical performance" refers to the act of performing a musical piece using instruments and voices.
[0531] "Extracting musical features" refers to the act of analyzing attributes such as pitch, intensity, and tempo of sound from audio data and extracting features.
[0532] The system of the present invention uses video and audio data analysis technology to provide fair and consistent artistic evaluation in real time. The system is mainly composed of three elements: a server, a terminal, and a user.
[0533] server
[0534] Data reception: The server first receives video and audio data from events such as figure skating and piano competitions. Specifically, it acquires streaming data via cameras and microphones using RTSP (Real-Time Streaming Protocol) or similar.
[0535] Data preprocessing: The server then preprocesses the received data. For video data, motion capture and image recognition technologies are used to extract human poses and movement characteristics. This is done using libraries such as OpenPose. For audio data, FFT (Fast Fourier Transform) is used to extract musical features such as pitch, intensity, and tempo.
[0536] Data Analysis: The preprocessed data is then fed into a generative AI model. The server uses deep learning models such as TensorFlow or PyTorch to analyze this data and calculate artistic scores. In the case of figure skating, scores are generated based on factors such as jump height, number of rotations, and spin speed. In the case of piano playing, scores are evaluated based on pitch accuracy, tempo consistency, and expressiveness.
[0537] Result storage: The analysis results (art scores) are stored in a database by the server, allowing for later analysis and review. This can be done using a SQL or NoSQL database.
[0538] Terminal
[0539] Display of results: The device displays the analysis results obtained from the server in real time. An interface is provided to visually present the results to competitors and judges. Specifically, the results are displayed using a web browser or a mobile app.
[0540] Accepting user input: The terminal also provides an interface for accepting user input, allowing operations such as starting and stopping a competition and checking results.
[0541] User
[0542] Hosting an event: A user hosts an event such as a figure skating or piano competition. They distribute entry forms to recruit participants, schedule the competition, and notify them.
[0543] Specific examples
[0544] In the case of figure skating
[0545] 1. A user hosts a figure skating competition.
[0546] Example: A user posts a competition and gathers participants online.
[0547] 2. The server receives live footage of the competition and collects data on jumps and spins.
[0548] Specific operation: The server receives video streaming data from the camera and extracts motion features using OpenPose.
[0549] 3. The server analyzes the preprocessed data using a generative AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[0550] Example prompt: "Analyze video of a figure skating competition and measure the jump height, number of rotations, and spin speed. Calculate an artistic score based on this data."
[0551] 4. The device displays the analysis results in real time and provides them to the competitors and judges.
[0552] Example: Results are displayed instantly on large screens at the stadium.
[0553] For piano performances
[0554] 1. A user hosts a piano competition.
[0555] Specific actions: Recruit participants using an entry form.
[0556] 2. The server receives the audio of the competition performance.
[0557] Example: A server captures audio data from a microphone in real time.
[0558] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[0559] Example: Extracting sound characteristics using FFT.
[0560] 4. The server analyzes the preprocessed data using a generative AI model to evaluate the sound characteristics and technical accuracy of the performance.
[0561] 5. The device displays the analysis results in real time and provides them to the user and judges.
[0562] Example: Analysis results are displayed instantly on smartphone apps and web apps.
[0563] In this way, the system of the present invention is able to provide fair and consistent grading for artistic competitions and reduce variability in grading.
[0564] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0565] Step 1: Collect data
[0566] A user hosts an event such as a figure skating or piano competition. Specifically, the user prepares an entry form, recruits participants, and sets and notifies the schedule of the competition. This completes the preparations for the competition.
[0567] The server receives video and audio data from the event in real time. Specific inputs include streaming video and audio data captured using a camera or microphone. The raw data is then stored on the server.
[0568] Step 2: Preprocessing the data
[0569] The server preprocesses the video and audio data it receives. Specifically, the server receives raw data and converts it into a format that is easy to analyze.
[0570] For video data, the server uses motion capture and image recognition technology to extract poses and movement characteristics of people, and specific outputs, such as jump height and spin rotations, are extracted using libraries such as OpenPose.
[0571] For audio data, the server uses FFT (Fast Fourier Transform) to extract musical features such as pitch, intensity, and tempo. Specific outputs include pitch accuracy, intensity distribution, and tempo fluctuation data.
[0572] Step 3: Analyze the data
[0573] The server inputs the preprocessed data into a generative AI model to calculate the artistic score. Specific inputs include preprocessed motion data and musical features. This is then input into an AI model such as TensorFlow or PyTorch.
[0574] In the case of figure skating, the server evaluates the height of jumps, number of rotations, speed of spins, etc. The output is a score for each element and an overall artistic score.
[0575] For piano performances, the server evaluates pitch accuracy, tempo consistency, and expressiveness, producing an output that measures the overall performance's technical accuracy and artistic evaluation score.
[0576] Step 4: Save the results
[0577] The server stores the analysis results (art scores) in a database. Specific inputs include the generated art scores and scores for each evaluation item. These are recorded in an SQL or NoSQL database.
[0578] As an output, the analysis results are stored in a format that can be used for later analysis and review, for example to analyze historical and trending results.
[0579] Step 5: View the results
[0580] The device displays the analysis results (art scores) in real time. Specific inputs include analysis results obtained from the server, which are then displayed on the user interface of a web browser or mobile app.
[0581] The output allows athletes and judges to check the evaluation results in real time. For example, the results may be displayed on a large screen at the competition venue or instantly on a smartphone app.
[0582] Step 6: User feedback
[0583] The user checks the evaluation results through the terminal and provides feedback. Specific inputs include feedback data from the user, which is sent to the server and stored.
[0584] As an output, the feedback is used to refine the system and adjust the evaluation criteria for the next iteration, thereby improving the overall accuracy of the system and the user experience.
[0585] (Application example 1)
[0586] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0587] With conventional methods, evaluation of live performances in physical venues is often subjective, making it difficult to achieve fair and consistent evaluations. Furthermore, the delay in real-time evaluation feedback often makes performers and audiences feel that the evaluations lack transparency and immediacy.
[0588] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0589] In this invention, the server includes means for receiving video data or audio data, means for pre-processing the received video data or audio data, means for analyzing the pre-processed data and calculating an artistic score, means for displaying the calculated artistic score, and means for capturing video and audio data of live performances through cameras and microphones in the store, thereby enabling fair and consistent real-time evaluation of live performances in physical stores.
[0590] "Video data" refers to dynamic or static image information acquired through a visual device such as a camera.
[0591] "Audio data" refers to an audio signal obtained through an audio device such as a microphone.
[0592] "Preprocessing" refers to a series of data manipulation steps that convert the received raw data into a format that is easier to analyze.
[0593] "Art score calculation" refers to the process of analyzing pre-processed data and generating a score based on specific evaluation criteria.
[0594] "Display means" refers to a display device or software interface that visually presents the analysis results to the user.
[0595] "In-store cameras" refer to video capture devices installed in physical stores to capture video data of live performances.
[0596] "In-store microphones" refer to audio capture devices installed in physical stores to capture audio data from live performances.
[0597] The system embodying this invention evaluates live performance in a physical store in real time and provides fair and consistent scoring. The system mainly uses a server, in-store cameras and microphones, a smartphone app, and a display.
[0598] Hardware and Software Configuration
[0599] 1. Server:
[0600] Role: Receiving, preprocessing, analyzing and storing data.
[0601] Software used:
[0602] Motion capture technology: OpenPose, etc.
[0603] Audio signal processing technology: librosa, etc.
[0604] AI analysis model: TensorFlow or PyTorch
[0605] 2. In-store cameras:
[0606] Role: Capture video data of the performance.
[0607] Example: High-definition video camera
[0608] 3. In-store microphones:
[0609] Role: Captures audio data of the performance.
[0610] Example: High-sensitivity microphone
[0611] 4. Smartphone App:
[0612] Role: Displays the score and accepts user input.
[0613] Example: Android or iOS app
[0614] 5. Display:
[0615] Role: Displaying evaluation results in-store.
[0616] Example: Large displays in stores
[0617] A natural language description of the program's operation
[0618] The server captures video data of the performance in real time through in-store cameras, and simultaneously captures audio data through in-store microphones. Motion capture technology is used to extract the poses and movement characteristics of the performers from the received video data, and audio signal processing technology is used to extract musical features such as pitch, intensity, and tempo from the audio data. The extracted features are input into an AI analysis model to calculate the performance's artistic score. The calculated artistic score is saved on the server and displayed in real time on a smartphone app and on in-store displays.
[0619] Specific examples
[0620] For example, imagine a live music concert is being held in a store. When the live performance begins, cameras in the store collect video data and microphones collect audio data, which are then sent to a server. The server preprocesses the received video data and analyzes the performers' movement characteristics using OpenPose. Musical features are extracted from the audio data using librosa, and these are input into an AI model built with TensorFlow and PyTorch. The model generates a score based on evaluation criteria, and the score is displayed in real time on a smartphone app or display.
[0621] Prompt Sentence Examples
[0622] Please use an AI model to analyze the beauty of jumps and spins in figure skating performance footage and provide a real-time evaluation score. Please enter the video data below. Please display the analysis results.
[0623] <Video data>
[0624] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0625] Step 1:
[0626] Data collection
[0627] When a live event begins, cameras and microphones in the store capture video and audio data of the performance in real time. These data are sent to a server. The cameras provide video data (e.g., the performers' movements) and the microphones provide audio data (e.g., the musical performance).
[0628] Input: Live performance video and audio data
[0629] Output: Video and audio data sent to the server
[0630] Step 2:
[0631] Data Preprocessing
[0632] The server preprocesses the received video data using OpenPose to extract human poses and movement features, and preprocesses the audio data using librosa to extract musical features such as pitch, intensity, and tempo, resulting in an analyzable dataset.
[0633] Input: Video and audio data sent to the server
[0634] Output: Preprocessed motion feature data and musical feature data
[0635] Step 3:
[0636] Data analysis
[0637] The server inputs the preprocessed data into an AI analysis model (e.g., using TensorFlow or PyTorch) to calculate the artistic score of the performance. The analysis model uses features such as the accuracy of the movements, the rhythm of the music, and the tempo as evaluation criteria.
[0638] Input: Preprocessed motion feature data and musical feature data
[0639] Output: Calculated art score
[0640] Step 4:
[0641] Saving the results
[0642] The server stores the calculated artistic scores in a database, which can be used for later review and analysis.
[0643] Input: Calculated art score
[0644] Output: Art points stored in the database
[0645] Step 5:
[0646] Displaying the results
[0647] The device displays the analysis results in real time, i.e., artistic scores, on a smartphone app and on displays in the store, allowing performers and audience members to immediately check the evaluation results.
[0648] Input: Art points stored in the database
[0649] Output: Smartphone app and artwork displayed on the screen
[0650] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0651] The system of the present invention includes a series of processes from receiving video or audio data to pre-processing, analysis, and display, and is combined with an emotion engine that recognizes the user's emotions. This system not only enables fair and consistent evaluation in artistic competitions and performances, but also takes into account the user's emotional responses.
[0652] System Overview
[0653] This system consists of a server, a terminal, an emotion engine, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user. The emotion engine recognizes the user's emotions and fine-tunes the evaluation based on their emotional response.
[0654] 1. Data collection
[0655] Users host figure skating and piano competitions.
[0656] The server receives video data (e.g., video of a figure skating performance) and audio data (e.g., audio of a piano performance) from the event in real time.
[0657] System operation details
[0658] 1. Data Collection
[0659] The server receives the video and audio data in real time.
[0660] In the case of video data, the data is obtained from a camera or video input source.
[0661] For audio data, data is obtained from a microphone or audio input source.
[0662] 2. Data Preprocessing
[0663] The server preprocesses the received video and audio data.
[0664] In the case of video data, motion capture and image recognition technology are used to extract a person's pose and movement characteristics.
[0665] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0666] 3. Data Analysis
[0667] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[0668] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[0669] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[0670] The analysis results (art scores) obtained are stored in a database.
[0671] 4. Emotion recognition
[0672] An emotion engine analyzes the facial expressions, vocal or body language of athletes, performers and spectators to identify their emotional state.
[0673] The server uses the output of the emotion engine to fine-tune the artistic score.
[0674] 5. Providing results
[0675] The terminal obtains the analysis results (art scores) in real time.
[0676] The terminal displays the analysis results to the competitor and judges.
[0677] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[0678] The server records the analysis results and stores them in a database for future analysis and review.
[0679] Specific examples
[0680] figure skating
[0681] 1. A user hosts a figure skating competition.
[0682] 2. The server receives live footage of the competition and collects data on jumps and spins.
[0683] 3. The server preprocesses the collected video data to extract information such as jump height and spin rotations.
[0684] 4. The server analyzes the preprocessed data using an AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[0685] 5. The emotion engine analyzes the facial expressions of the audience and judges and measures the level of emotion based on each element.
[0686] 6. The server fine-tunes the artistic score based on the output of the emotion engine.
[0687] 7. The device displays the analysis results in real time and provides them to the competitors and judges.
[0688] Piano performance
[0689] 1. A user hosts a piano competition.
[0690] 2. The server receives the audio of the competition performance.
[0691] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[0692] 4. The server analyzes the preprocessed data using an AI model to evaluate the sound characteristics and technical accuracy of the performance.
[0693] 5. The emotion engine analyzes the facial expressions of the audience and judges and measures the level of emotion based on each element.
[0694] 6. The server fine-tunes the artistic score based on the output of the emotion engine.
[0695] 7. The device displays the analysis results in real time and provides them to the user and judges.
[0696] In this way, each step works together to enable the system of the present invention to provide fair and consistent evaluations, and also to take into account the user's emotional responses.
[0697] The processing flow will be explained below.
[0698] Processing flow and each processing step
[0699] Step 1: Receiving Data
[0700] A user hosts a figure skating competition or a piano competition.
[0701] A server receives video or audio data from an event in real time.
[0702] In the case of video data, the data is obtained from a camera or video input source.
[0703] For audio data, data is obtained from a microphone or audio input source.
[0704] Step 2: Save your data
[0705] Create a directory for the server to temporarily store the data it receives.
[0706] The server temporarily saves the received data stream in the "raw_data" directory.
[0707] Step 3: Preprocessing the data
[0708] The server reads the saved data from the "raw_data" directory and performs preprocessing.
[0709] In the case of video data, motion capture and image recognition technology are used to extract people's poses and movement characteristics.
[0710] Motion capture is a technology that identifies the joint positions of a person in a video.
[0711] Image recognition technology is a technology that recognizes specific movements (e.g., jumps and spins).
[0712] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0713] Pitch is a feature that indicates the pitch of a sound.
[0714] Intensity is a feature that indicates the loudness of a sound.
[0715] Tempo is a characteristic that indicates the speed of a performance.
[0716] The server saves the preprocessed data in the "processed_data" directory.
[0717] Step 4: Analyze the data
[0718] The server reads the preprocessed data from the "processed_data" directory.
[0719] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[0720] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[0721] The jump height refers to the height from the start point of the jump to the landing point.
[0722] The number of rotations refers to the number of rotations during the jump.
[0723] The rotation speed of the spin indicates the speed of rotation during the spin.
[0724] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[0725] The accuracy of the sound is evaluated by whether the sound is produced according to the musical score.
[0726] Expressiveness is measured by the performer's richness of emotion and expression.
[0727] Rhythmic consistency is measured by whether the performance maintains a consistent tempo.
[0728] The server stores the analysis results (art scores) in a database.
[0729] Step 5: Emotion Recognition
[0730] The emotion engine analyzes the facial expressions, voice or body language of athletes and spectators to identify their emotional state.
[0731] Facial expression analysis is a technology that analyzes facial expressions from camera footage and identifies emotions.
[0732] Voice analysis is a technology that analyzes the tone and strength of a voice to identify emotions.
[0733] Body language analysis is a technique that analyzes gestures and postures to identify emotions.
[0734] The server uses the output of the emotion engine to fine-tune the artistic score.
[0735] The emotional score is a numerical representation of the user's emotional response.
[0736] To fine-tune the artistic score, the emotional score is taken into account in the final evaluation.
[0737] Step 6: Delivering results
[0738] The terminal obtains the analysis results (art scores) in real time.
[0739] The terminal provides a display device for displaying the analysis results to the competitors and judges.
[0740] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[0741] The server records the analysis results and stores them in a database for future analysis and review.
[0742] In this way, each step works together to enable the system of the present invention to provide fair and consistent evaluations, and also to take into account the user's emotional responses.
[0743] Example 2
[0744] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0745] Conventional evaluation systems for artistic competitions and performances have difficulty providing fair and consistent evaluations. In particular, they lack systems that reflect the emotional reactions of audiences and judges in their evaluations, resulting in a lack of objectivity in the evaluations. Furthermore, they lack the ability to provide evaluations in real time and to store evaluation results.
[0746] The identification processing by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving video data or audio data, means for preprocessing the received video data or audio data, means for analyzing the preprocessed data and calculating an artistic score, means for recognizing the emotions of the contestants and spectators, means for fine-tuning the artistic score based on the emotion recognition results, and means for displaying the calculated artistic score. This makes it possible to provide fair and consistent evaluations and realize evaluations that take into account the emotional reactions of spectators and judges.
[0747] "Video data" refers to data that includes image information obtained from a camera, video input source, or the like.
[0748] "Audio data" refers to data containing sound information obtained from a microphone, audio input source, or the like.
[0749] "Preprocessing" refers to performing initial processing to extract specific features from received video and audio data.
[0750] "Analysis" refers to the evaluation of pre-processed data using AI models or other means to produce a specific result (e.g., artistic score).
[0751] "Artistic Score" is a score calculated based on the received video and audio data, or the emotional response of the competitors and spectators.
[0752] "Emotion recognition" is the process of analyzing the facial expressions, voice, body language, etc. of athletes and spectators to identify their emotional state.
[0753] "Fine-tuning based on emotion recognition results" means adjusting the aforementioned analysis results (artistic score) based on the data obtained through emotion recognition.
[0754] "Display" refers to the visual presentation of the analyzed results and fine-tuned artistic points.
[0755] The system of the present invention realizes a series of processes from receiving video or audio data, to preprocessing, analysis, display, and further fine-tuning evaluation by emotion recognition. This section explains the specific implementation method.
[0756] This system mainly consists of a server, a terminal, an emotion engine, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user. The emotion engine recognizes the user's emotions and fine-tunes the evaluation based on their emotional response.
[0757] Data collection
[0758] First, a user hosts an artistic competition or performance event (e.g., figure skating or piano competition). The server receives real-time video and audio data from the event. Video data comes from cameras and other video input sources, and audio data comes from microphones and other audio input sources.
[0759] Data Preprocessing
[0760] The server preprocesses the received video and audio data. For video data, motion capture and image recognition technology are used to extract the poses and movement characteristics of the person. For audio data, audio signal processing technology is used to extract musical features such as the pitch, intensity, and tempo of the sound. For example, the height of a figure skater's jump or the number of rotations in a spin, or the accuracy and rhythmic consistency of a piano performance can be obtained.
[0761] Data analysis
[0762] The server inputs the preprocessed data into an AI model to calculate artistic scores. In the case of video data, the AI model evaluates jump height, number of rotations, spin speed, etc., and in the case of audio data, it evaluates sound accuracy, expressiveness, rhythmic consistency, etc. The analysis results (artistic scores) obtained are stored in a database.
[0763] emotion recognition
[0764] The emotion engine analyzes the facial expressions, voice, and body language of competitors, performers, and audience members to identify their emotional state. The server uses the output of the emotion engine to fine-tune the artistic score, enabling a fair evaluation that reflects not only pure technical evaluation but also the emotional reactions of the audience and judges.
[0765] Providing results
[0766] The device acquires the analysis results (artistic scores) in real time and displays them to the competitors and judges. The user can check the displayed analysis results and use them to evaluate the competition and performance. The server also records the analysis results and stores them in a database for future analysis and review.
[0767] Specific examples
[0768] In the case of figure skating: A user hosts a competition, and the server receives live footage of the competition and collects data on jumps and spins. The server preprocesses this data, analyzes the jump height and spin rate using an AI model, and calculates and adjusts the artistic score. The results are displayed on the device in real time and provided to competitors and judges.
[0769] In the case of piano performance: A user hosts a piano competition, and the server receives the audio of the competition performance. The server preprocesses the audio data and analyzes pitch, intensity, tempo, etc. using an AI model to evaluate the accuracy of the sound and the technical accuracy of the performance. The emotion engine analyzes the facial expressions of the audience and judges and fine-tunes the artistic score based on their emotions. The results are displayed on the device in real time and provided to the user and judges.
[0770] Example prompt sentence:
[0771] Explain the processing steps of a system that evaluates the height of figure skating jumps and the number of rotations in spins, and adjusts the final artistic score taking into account the emotional response of the audience.
[0772]
[0773] The system analyzes musical characteristics such as pitch, intensity, and tempo during piano performances, and fine-tunes its evaluation based on the emotional responses of the audience and judges.
[0774] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0775] Step 1: Collect data
[0776] A user hosts events such as figure skating and piano competitions.
[0777] The server receives video or audio data from the event in real time. Specifically, video data is obtained from a camera or video input, and audio data is obtained from a microphone or audio input.
[0778] Input: Video data from a camera or video input, audio data from a microphone or audio input
[0779] Output: Raw data stored on the server
[0780] Step 2: Data Preprocessing
[0781] The server preprocesses the video and audio data received.
[0782] In the case of video data, motion capture and image recognition technology are used to extract a person's pose and movement characteristics.
[0783] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0784] Input: Received raw video and audio data
[0785] Output: Preprocessed feature data (e.g., person pose data, sound characteristic data)
[0786] Step 3: Data analysis
[0787] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[0788] When using video data, the AI model evaluates factors such as jump height, number of rotations, and spin speed.
[0789] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[0790] Input: Preprocessed feature data
[0791] Output: Art score calculated by the AI model
[0792] Step 4: Emotion Recognition
[0793] The emotion engine analyzes the facial expressions, voice or body language of athletes, performers and spectators to identify their emotional state.
[0794] The server fine-tunes the artistic score based on the output of the emotion engine.
[0795] Input: Emotional data obtained from the camera or microphone (facial expressions, voice, body language)
[0796] Output: Emotional state analyzed by emotion engine, fine-tuned artistic score
[0797] Step 5: Delivering results
[0798] The terminal obtains the analysis results (art scores) in real time.
[0799] The terminal displays the analysis results to the competitors and judges.
[0800] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[0801] The server records the analysis results and stores them in a database for future analysis and review.
[0802] Input: Fine-tuned art scores, emotion recognition results
[0803] Output: Final evaluation results displayed in real time on the device, evaluation results stored in the database
[0804] (Application example 2)
[0805] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0806] Currently, there are almost no systems in physical stores that can grasp customer emotions and behavior in real time and provide appropriate responses based on that information. This makes it difficult to provide services that immediately reflect customer needs and emotions, limiting the improvement of customer satisfaction. It is also difficult to introduce a fair and consistent evaluation system.
[0807] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving video data or audio data, means for preprocessing the received video data or audio data, means for analyzing the preprocessed data and calculating an artistic score, means for displaying the calculated artistic score, and means for recognizing the user's emotions and fine-tuning the evaluation based on the emotional response. This makes it possible to analyze customer emotions and behavior in real time in physical stores and provide appropriate feedback to store staff immediately. Furthermore, using an emotion recognition engine enables a fair and consistent evaluation system and improves customer satisfaction.
[0808] "Video data" refers to image or video data acquired through a camera, video device, or the like.
[0809] "Audio data" refers to sound or voice data acquired through a microphone, recording device, or the like.
[0810] "Receiving means" refers to a function or system that allows a server or terminal to receive video data or audio data from the outside.
[0811] "Preprocessing means" refers to a function or system that processes received video data and audio data and converts them into a format suitable for analysis.
[0812] "Analysis means" refers to a function or system for calculating evaluation points based on preprocessed data.
[0813] "Artistic score" is a numerical evaluation of video and audio, and is an index used to evaluate technical and artistic aspects.
[0814] "Display means" refers to a function or system for visually displaying the calculated evaluation scores and analysis results.
[0815] "Emotion recognition" is a technology that identifies a user's emotional state based on data such as facial expressions and voice.
[0816] "Fine-tuning of the evaluation" is the process of adjusting the evaluation score calculated based on the emotion recognition results.
[0817] A "camera" is a photographing device for acquiring video data.
[0818] A "microphone" is a recording device for capturing audio data.
[0819] An "emotion engine" is software or hardware for analyzing a user's emotional state.
[0820] A "physical store" is a sales location set up in a physical location where customers visit in person to make purchases.
[0821] A "store clerk" is an employee who provides services to customers in a physical store.
[0822] The system of the present invention recognizes customer emotions and provides feedback to store clerks in real time. This system consists of a server, a camera, a microphone, a terminal, and an emotion engine. Specific hardware and software used include OpenCV, Keras, librosa, etc.
[0823] The system is configured as follows:
[0824] 1. Data Collection:
[0825] The server receives video and audio data from cameras and microphones installed in the store. The cameras capture customers' facial expressions in real time, and the microphones record their voices.
[0826] 2. Data preprocessing:
[0827] The received video and audio data are preprocessed by the server. The video data undergoes face recognition and facial expression analysis using OpenCV. The audio data undergoes audio feature extraction (e.g., MFCC) using librosa.
[0828] 3. Data Analysis:
[0829] The pre-processed data is fed into a generative AI model using Keras to analyze the customer's emotional state. Emotions are recognized from facial and vocal characteristics, and specific emotional responses (e.g., joy, anger, sadness, etc.) are identified by an emotion engine.
[0830] 4. Providing results:
[0831] The server displays the analysis results on the terminal in real time. The store clerk can then use this information to respond appropriately to the customer. For example, if the customer is angry, a more careful response appropriate to the situation will be recommended. If the customer is satisfied, the usual service will be provided.
[0832] As a concrete example, we will show an example of inputting a prompt sentence into a generative AI model.
[0833] Prompt: "Implement a system in a brick-and-mortar store that recognizes customer emotions and instructs store associates on appropriate responses based on facial and voice data. Key scenarios to consider include when a customer is angry or sad."
[0834] The above configuration and procedures make it possible to improve the customer experience in physical stores, and are expected to increase customer satisfaction.
[0835] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0836] Step 1: Collect data
[0837] The server receives video and audio data of customers in the store via a camera and microphone. Specifically, the camera captures the customer's face and body movements, and the microphone records the customer's voice. The input is video data from the camera and audio data from the microphone, which are then sent to the server.
[0838] Step 2: Preprocessing the data
[0839] The server preprocesses the received video and audio data. Specifically, it uses OpenCV to recognize faces and facial expressions from the video data and normalizes the images. For audio data, it uses librosa to convert the audio signal into features (e.g., MFCCs). The input is raw video and audio data, and the output is preprocessed face image data and audio feature data.
[0840] Step 3: Analyze the data
[0841] The server analyzes the preprocessed data. Specifically, a generative AI model using Keras is used to estimate the emotional state from the facial image data and the emotional state of the voice from the voice feature data. The input is the preprocessed facial image data and voice feature data, and the output is the customer's emotional state (e.g., joy, anger, sadness, etc.).
[0842] Step 4: Applying the Emotion Engine
[0843] The server further processes the analysis results using an emotion engine to refine the accuracy. The emotion engine refines the overall emotion rating based on the analyzed emotional state. The input is the analyzed emotional state, and the output is a fine-tuned emotion rating.
[0844] Step 5: View the results
[0845] The server displays the adjusted emotion evaluation on the terminal in real time. Specifically, it notifies the store clerk of the analysis results and the fine-tuned emotion evaluation. The input is the fine-tuned emotion evaluation, and the output is a notification message displayed on the store clerk's terminal. For example, if the customer is angry, the message displayed is "The customer is angry. Please be careful how you respond."
[0846] Through these steps, the system can recognize customer emotions in real time in physical stores and provide immediate and appropriate feedback to store staff.
[0847] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0848] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0849] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0850] [Third embodiment]
[0851] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0852] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0853] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0854] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0855] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0856] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0857] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0858] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0859] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0860] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0861] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0862] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0863] The system of the present invention analyzes video or audio data in real time to provide fair and consistent evaluation in artistic scoring competitions. The system has a series of functions for receiving, preprocessing, and analyzing video or audio data, and displaying the resulting artistic scores. The specific operation of the system is described below.
[0864] System Overview
[0865] This system consists of a server, a terminal, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user.
[0866] 1. Data collection
[0867] Users host events such as figure skating and piano competitions.
[0868] A server receives video data (e.g., video of a figure skating performance) and audio data (e.g., audio of a piano performance) from an event in real time.
[0869] System operation details
[0870] 1. Data Collection
[0871] A server receives a live stream of a figure skating or piano competition.
[0872] In the case of video data, the server acquires data in real time from cameras and other video input sources.
[0873] For audio data, the server captures data in real time from a microphone or other audio input source.
[0874] 2. Data Preprocessing
[0875] The server converts the received video and audio data into an easy-to-understand format.
[0876] In the case of figure skating, motion capture and image recognition technology are used to extract a person's poses and movement characteristics.
[0877] In the case of piano performance, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0878] 3. Data Analysis
[0879] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[0880] In figure skating, the height of the jump, the number of rotations, and the speed of the spins are evaluated.
[0881] Piano performance is assessed on accuracy of tone, expressiveness, and rhythmic consistency.
[0882] The analyzed results (art scores) are stored in a database.
[0883] 4. Providing results
[0884] The terminal displays the analysis results (art scores) in real time.
[0885] The terminal is equipped with a display device for displaying information to competitors and judges.
[0886] The server records the results and makes them available for later analysis and review.
[0887] Specific examples
[0888] figure skating
[0889] 1. A user hosts a figure skating competition.
[0890] 2. The server receives live footage of the competition and collects data on jumps and spins.
[0891] 3. The server preprocesses the collected video data to extract information such as jump height and spin rotations.
[0892] 4. The server analyzes the preprocessed data using an AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[0893] 5. The device displays the analysis results in real time and provides them to the competitors and judges.
[0894] Piano performance
[0895] 1. A user hosts a piano competition.
[0896] 2. The server receives the audio of the competition performance.
[0897] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[0898] 4. The server analyzes the preprocessed data using an AI model to evaluate the sound characteristics and technical accuracy of the performance.
[0899] 5. The device displays the analysis results in real time and provides them to the user and judges.
[0900] In this way, the system of the present invention is able to provide fair and consistent grading for artistic competitions and reduce variability in grading.
[0901] The processing flow will be explained below.
[0902] Processing flow and each processing step
[0903] Step 1: Receiving Data
[0904] A user hosts a figure skating competition or a piano competition.
[0905] A server receives video or audio data from an event in real time.
[0906] In the case of video data, the data is obtained from a camera or video input source.
[0907] For audio data, data is obtained from a microphone or audio input source.
[0908] Step 2: Save your data
[0909] Create a directory for the server to temporarily store the data it receives.
[0910] The server temporarily saves the received data stream in the "raw_data" directory.
[0911] Step 3: Preprocessing the data
[0912] The server reads the saved data from the "raw_data" directory.
[0913] In the case of video data, motion capture and image recognition technology are used to extract people's poses and movement characteristics.
[0914] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[0915] The server saves the preprocessed data in the "processed_data" directory.
[0916] Step 4: Analyze the data
[0917] The server reads the preprocessed data from the "processed_data" directory.
[0918] The server inputs the preprocessed data into an AI model to calculate the artistic score.
[0919] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[0920] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[0921] The server stores the analysis results (art scores) in a database.
[0922] Step 5: Delivering results
[0923] The terminal obtains the analysis results (art scores) in real time.
[0924] The terminal provides a display device for displaying the analysis results to the competitors and judges.
[0925] The user can check the displayed analysis results and use them to evaluate the competition or performance.
[0926] The server records the analysis results and stores them in a database for future analysis and review.
[0927] In this way, with each step working together, the system of the present invention provides fair and consistent evaluation of artistic competitions and performances.
[0928] Example 1
[0929] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0930] Conventional evaluation methods for artistic competitions often rely on the subjective judgment and experience of judges, resulting in a lack of fairness and consistency. Real-time evaluation is also difficult, resulting in variations and delays in evaluation. The present invention aims to solve these problems and provide a system that provides fair and consistent evaluation in real time.
[0931] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0932] In this invention, the server includes a means for receiving video data or audio data, a means for preprocessing the received video data or audio data, and a means for inputting the preprocessed data into a generative AI model to calculate artistic scores, thereby enabling the server to analyze the video data or audio data in real time and provide fair and consistent evaluations.
[0933] "Video Data" refers to visual information of an action or scene captured from a camera or other video input device.
[0934] "Audio Data" refers to sound information collected from a microphone or other audio input device.
[0935] "Means for receiving" refers to the technical means by which the server retrieves video and audio data from a network or device.
[0936] "Pre-processing means" refers to technical means for converting received video and audio data into an understandable format. Examples include motion capture and audio signal processing techniques.
[0937] A "generative AI model" is a model that uses artificial intelligence technology to analyze data and generate scores or ratings for specific purposes.
[0938] "Means for calculating artistic scores" refers to the technical means for inputting preprocessed data into a generative AI model and conducting a quantitative evaluation.
[0939] "Means for displaying" refers to the technical means for visually presenting the artistic score, which is the result of the analysis, to users and judges.
[0940] "Personal action" refers to a figure skating routine or other artistic movement performed by a particular person.
[0941] "Extracting pose and movement features" refers to the act of analyzing specific postures and movements from video data and extracting related features.
[0942] "Musical performance" refers to the act of performing a musical piece using instruments and voices.
[0943] "Extracting musical features" refers to the act of analyzing attributes such as pitch, intensity, and tempo of sound from audio data and extracting features.
[0944] The system of the present invention uses video and audio data analysis technology to provide fair and consistent artistic evaluation in real time. The system is mainly composed of three elements: a server, a terminal, and a user.
[0945] server
[0946] Data reception: The server first receives video and audio data from events such as figure skating and piano competitions. Specifically, it acquires streaming data via cameras and microphones using RTSP (Real-Time Streaming Protocol) or similar.
[0947] Data preprocessing: The server then preprocesses the received data. For video data, motion capture and image recognition technologies are used to extract human poses and movement characteristics. This is done using libraries such as OpenPose. For audio data, FFT (Fast Fourier Transform) is used to extract musical features such as pitch, intensity, and tempo.
[0948] Data Analysis: The preprocessed data is then fed into a generative AI model. The server uses deep learning models such as TensorFlow or PyTorch to analyze this data and calculate artistic scores. In the case of figure skating, scores are generated based on factors such as jump height, number of rotations, and spin speed. In the case of piano playing, scores are evaluated based on pitch accuracy, tempo consistency, and expressiveness.
[0949] Result storage: The analysis results (art scores) are stored in a database by the server, allowing for later analysis and review. This can be done using a SQL or NoSQL database.
[0950] Terminal
[0951] Display of results: The device displays the analysis results obtained from the server in real time. An interface is provided to visually present the results to competitors and judges. Specifically, the results are displayed using a web browser or a mobile app.
[0952] Accepting user input: The terminal also provides an interface for accepting user input, allowing operations such as starting and stopping a competition and checking results.
[0953] User
[0954] Hosting an event: A user hosts an event such as a figure skating or piano competition. They distribute entry forms to recruit participants, schedule the competition, and notify them.
[0955] Specific examples
[0956] In the case of figure skating
[0957] 1. A user hosts a figure skating competition.
[0958] Example: A user posts a competition and gathers participants online.
[0959] 2. The server receives live footage of the competition and collects data on jumps and spins.
[0960] Specific operation: The server receives video streaming data from the camera and extracts motion features using OpenPose.
[0961] 3. The server analyzes the preprocessed data using a generative AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[0962] Example prompt: "Analyze video of a figure skating competition and measure the jump height, number of rotations, and spin speed. Calculate an artistic score based on this data."
[0963] 4. The device displays the analysis results in real time and provides them to the competitors and judges.
[0964] Example: Results are displayed instantly on large screens at the stadium.
[0965] For piano performances
[0966] 1. A user hosts a piano competition.
[0967] Specific actions: Recruit participants using an entry form.
[0968] 2. The server receives the audio of the competition performance.
[0969] Example: A server captures audio data from a microphone in real time.
[0970] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[0971] Example: Extracting sound characteristics using FFT.
[0972] 4. The server analyzes the preprocessed data using a generative AI model to evaluate the sound characteristics and technical accuracy of the performance.
[0973] 5. The device displays the analysis results in real time and provides them to the user and judges.
[0974] Example: Analysis results are displayed instantly on smartphone apps and web apps.
[0975] In this way, the system of the present invention is able to provide fair and consistent grading for artistic competitions and reduce variability in grading.
[0976] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0977] Step 1: Collect data
[0978] A user hosts an event such as a figure skating or piano competition. Specifically, the user prepares an entry form, recruits participants, and sets and notifies the schedule of the competition. This completes the preparations for the competition.
[0979] The server receives video and audio data from the event in real time. Specific inputs include streaming video and audio data captured using a camera or microphone. The raw data is then stored on the server.
[0980] Step 2: Preprocessing the data
[0981] The server preprocesses the video and audio data it receives. Specifically, the server receives raw data and converts it into a format that is easy to analyze.
[0982] For video data, the server uses motion capture and image recognition technology to extract poses and movement characteristics of people, and specific outputs, such as jump height and spin rotations, are extracted using libraries such as OpenPose.
[0983] For audio data, the server uses FFT (Fast Fourier Transform) to extract musical features such as pitch, intensity, and tempo. Specific outputs include pitch accuracy, intensity distribution, and tempo fluctuation data.
[0984] Step 3: Analyze the data
[0985] The server inputs the preprocessed data into a generative AI model to calculate the artistic score. Specific inputs include preprocessed motion data and musical features. This is then input into an AI model such as TensorFlow or PyTorch.
[0986] In the case of figure skating, the server evaluates the height of jumps, number of rotations, speed of spins, etc. The output is a score for each element and an overall artistic score.
[0987] For piano performances, the server evaluates pitch accuracy, tempo consistency, and expressiveness, producing an output that measures the overall performance's technical accuracy and artistic evaluation score.
[0988] Step 4: Save the results
[0989] The server stores the analysis results (art scores) in a database. Specific inputs include the generated art scores and scores for each evaluation item. These are recorded in an SQL or NoSQL database.
[0990] As an output, the analysis results are stored in a format that can be used for later analysis and review, for example to analyze historical and trending results.
[0991] Step 5: View the results
[0992] The device displays the analysis results (art scores) in real time. Specific inputs include analysis results obtained from the server, which are then displayed on the user interface of a web browser or mobile app.
[0993] The output allows athletes and judges to check the evaluation results in real time. For example, the results may be displayed on a large screen at the competition venue or instantly on a smartphone app.
[0994] Step 6: User feedback
[0995] The user checks the evaluation results through the terminal and provides feedback. Specific inputs include feedback data from the user, which is sent to the server and stored.
[0996] As an output, the feedback is used to refine the system and adjust the evaluation criteria for the next iteration, thereby improving the overall accuracy of the system and the user experience.
[0997] (Application example 1)
[0998] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0999] With conventional methods, evaluation of live performances in physical venues is often subjective, making it difficult to achieve fair and consistent evaluations. Furthermore, the delay in real-time evaluation feedback often makes performers and audiences feel that the evaluations lack transparency and immediacy.
[1000] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1001] In this invention, the server includes means for receiving video data or audio data, means for pre-processing the received video data or audio data, means for analyzing the pre-processed data and calculating an artistic score, means for displaying the calculated artistic score, and means for capturing video and audio data of live performances through cameras and microphones in the store, thereby enabling fair and consistent real-time evaluation of live performances in physical stores.
[1002] "Video data" refers to dynamic or static image information acquired through a visual device such as a camera.
[1003] "Audio data" refers to an audio signal obtained through an audio device such as a microphone.
[1004] "Preprocessing" refers to a series of data manipulation steps that convert the received raw data into a format that is easier to analyze.
[1005] "Art score calculation" refers to the process of analyzing pre-processed data and generating a score based on specific evaluation criteria.
[1006] "Display means" refers to a display device or software interface that visually presents the analysis results to the user.
[1007] "In-store cameras" refer to video capture devices installed in physical stores to capture video data of live performances.
[1008] "In-store microphones" refer to audio capture devices installed in physical stores to capture audio data from live performances.
[1009] The system embodying this invention evaluates live performance in a physical store in real time and provides fair and consistent scoring. The system mainly uses a server, in-store cameras and microphones, a smartphone app, and a display.
[1010] Hardware and Software Configuration
[1011] 1. Server:
[1012] Role: Receiving, preprocessing, analyzing and storing data.
[1013] Software used:
[1014] Motion capture technology: OpenPose, etc.
[1015] Audio signal processing technology: librosa, etc.
[1016] AI analysis model: TensorFlow or PyTorch
[1017] 2. In-store cameras:
[1018] Role: Capture video data of the performance.
[1019] Example: High-definition video camera
[1020] 3. In-store microphones:
[1021] Role: Captures audio data of the performance.
[1022] Example: High-sensitivity microphone
[1023] 4. Smartphone App:
[1024] Role: Displays the score and accepts user input.
[1025] Example: Android or iOS app
[1026] 5. Display:
[1027] Role: Displaying evaluation results in-store.
[1028] Example: Large displays in stores
[1029] A natural language description of the program's operation
[1030] The server captures video data of the performance in real time through in-store cameras, and simultaneously captures audio data through in-store microphones. Motion capture technology is used to extract the poses and movement characteristics of the performers from the received video data, and audio signal processing technology is used to extract musical features such as pitch, intensity, and tempo from the audio data. The extracted features are input into an AI analysis model to calculate the performance's artistic score. The calculated artistic score is saved on the server and displayed in real time on a smartphone app and on in-store displays.
[1031] Specific examples
[1032] For example, imagine a live music concert is being held in a store. When the live performance begins, cameras in the store collect video data and microphones collect audio data, which are then sent to a server. The server preprocesses the received video data and analyzes the performers' movement characteristics using OpenPose. Musical features are extracted from the audio data using librosa, and these are input into an AI model built with TensorFlow and PyTorch. The model generates a score based on evaluation criteria, and the score is displayed in real time on a smartphone app or display.
[1033] Prompt Sentence Examples
[1034] Please use an AI model to analyze the beauty of jumps and spins in figure skating performance footage and provide a real-time evaluation score. Please enter the video data below. Please display the analysis results.
[1035] <Video data>
[1036] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1037] Step 1:
[1038] Data collection
[1039] When a live event begins, cameras and microphones in the store capture video and audio data of the performance in real time. These data are sent to a server. The cameras provide video data (e.g., the performers' movements) and the microphones provide audio data (e.g., the musical performance).
[1040] Input: Live performance video and audio data
[1041] Output: Video and audio data sent to the server
[1042] Step 2:
[1043] Data Preprocessing
[1044] The server preprocesses the received video data using OpenPose to extract human poses and movement features, and preprocesses the audio data using librosa to extract musical features such as pitch, intensity, and tempo, resulting in an analyzable dataset.
[1045] Input: Video and audio data sent to the server
[1046] Output: Preprocessed motion feature data and musical feature data
[1047] Step 3:
[1048] Data analysis
[1049] The server inputs the preprocessed data into an AI analysis model (e.g., using TensorFlow or PyTorch) to calculate the artistic score of the performance. The analysis model uses features such as the accuracy of the movements, the rhythm of the music, and the tempo as evaluation criteria.
[1050] Input: Preprocessed motion feature data and musical feature data
[1051] Output: Calculated art score
[1052] Step 4:
[1053] Saving the results
[1054] The server stores the calculated artistic scores in a database, which can be used for later review and analysis.
[1055] Input: Calculated art score
[1056] Output: Art points stored in the database
[1057] Step 5:
[1058] Displaying the results
[1059] The device displays the analysis results in real time, i.e., artistic scores, on a smartphone app and on displays in the store, allowing performers and audience members to immediately check the evaluation results.
[1060] Input: Art points stored in the database
[1061] Output: Smartphone app and artwork displayed on the screen
[1062] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1063] The system of the present invention includes a series of processes from receiving video or audio data to pre-processing, analysis, and display, and is combined with an emotion engine that recognizes the user's emotions. This system not only enables fair and consistent evaluation in artistic competitions and performances, but also takes into account the user's emotional responses.
[1064] System Overview
[1065] This system consists of a server, a terminal, an emotion engine, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user. The emotion engine recognizes the user's emotions and fine-tunes the evaluation based on their emotional response.
[1066] 1. Data collection
[1067] Users host figure skating and piano competitions.
[1068] The server receives video data (e.g., video of a figure skating performance) and audio data (e.g., audio of a piano performance) from the event in real time.
[1069] System operation details
[1070] 1. Data Collection
[1071] The server receives the video and audio data in real time.
[1072] In the case of video data, the data is obtained from a camera or video input source.
[1073] For audio data, data is obtained from a microphone or audio input source.
[1074] 2. Data Preprocessing
[1075] The server preprocesses the received video and audio data.
[1076] In the case of video data, motion capture and image recognition technology are used to extract a person's pose and movement characteristics.
[1077] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[1078] 3. Data Analysis
[1079] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[1080] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[1081] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[1082] The analysis results (art scores) obtained are stored in a database.
[1083] 4. Emotion recognition
[1084] An emotion engine analyzes the facial expressions, vocal or body language of athletes, performers and spectators to identify their emotional state.
[1085] The server uses the output of the emotion engine to fine-tune the artistic score.
[1086] 5. Providing results
[1087] The terminal obtains the analysis results (art scores) in real time.
[1088] The terminal displays the analysis results to the competitor and judges.
[1089] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[1090] The server records the analysis results and stores them in a database for future analysis and review.
[1091] Specific examples
[1092] figure skating
[1093] 1. A user hosts a figure skating competition.
[1094] 2. The server receives live footage of the competition and collects data on jumps and spins.
[1095] 3. The server preprocesses the collected video data to extract information such as jump height and spin rotations.
[1096] 4. The server analyzes the preprocessed data using an AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[1097] 5. The emotion engine analyzes the facial expressions of the audience and judges and measures the level of emotion based on each element.
[1098] 6. The server fine-tunes the artistic score based on the output of the emotion engine.
[1099] 7. The device displays the analysis results in real time and provides them to the competitors and judges.
[1100] Piano performance
[1101] 1. A user hosts a piano competition.
[1102] 2. The server receives the audio of the competition performance.
[1103] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[1104] 4. The server analyzes the preprocessed data using an AI model to evaluate the sound characteristics and technical accuracy of the performance.
[1105] 5. The emotion engine analyzes the facial expressions of the audience and judges and measures the level of emotion based on each element.
[1106] 6. The server fine-tunes the artistic score based on the output of the emotion engine.
[1107] 7. The device displays the analysis results in real time and provides them to the user and judges.
[1108] In this way, each step works together to enable the system of the present invention to provide fair and consistent evaluations, and also to take into account the user's emotional responses.
[1109] The processing flow will be explained below.
[1110] Processing flow and each processing step
[1111] Step 1: Receiving Data
[1112] A user hosts a figure skating competition or a piano competition.
[1113] A server receives video or audio data from an event in real time.
[1114] In the case of video data, the data is obtained from a camera or video input source.
[1115] For audio data, data is obtained from a microphone or audio input source.
[1116] Step 2: Save your data
[1117] Create a directory for the server to temporarily store the data it receives.
[1118] The server temporarily saves the received data stream in the "raw_data" directory.
[1119] Step 3: Preprocessing the data
[1120] The server reads the saved data from the "raw_data" directory and performs preprocessing.
[1121] In the case of video data, motion capture and image recognition technology are used to extract people's poses and movement characteristics.
[1122] Motion capture is a technology that identifies the joint positions of a person in a video.
[1123] Image recognition technology is a technology that recognizes specific movements (e.g., jumps and spins).
[1124] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[1125] Pitch is a feature that indicates the pitch of a sound.
[1126] Intensity is a feature that indicates the loudness of a sound.
[1127] Tempo is a characteristic that indicates the speed of a performance.
[1128] The server saves the preprocessed data in the "processed_data" directory.
[1129] Step 4: Analyze the data
[1130] The server reads the preprocessed data from the "processed_data" directory.
[1131] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[1132] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[1133] The jump height refers to the height from the start point of the jump to the landing point.
[1134] The number of rotations refers to the number of rotations during the jump.
[1135] The rotation speed of the spin indicates the speed of rotation during the spin.
[1136] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[1137] The accuracy of the sound is evaluated by whether the sound is produced according to the musical score.
[1138] Expressiveness is measured by the performer's richness of emotion and expression.
[1139] Rhythmic consistency is measured by whether the performance maintains a consistent tempo.
[1140] The server stores the analysis results (art scores) in a database.
[1141] Step 5: Emotion Recognition
[1142] The emotion engine analyzes the facial expressions, voice or body language of athletes and spectators to identify their emotional state.
[1143] Facial expression analysis is a technology that analyzes facial expressions from camera footage and identifies emotions.
[1144] Voice analysis is a technology that analyzes the tone and strength of a voice to identify emotions.
[1145] Body language analysis is a technique that analyzes gestures and postures to identify emotions.
[1146] The server uses the output of the emotion engine to fine-tune the artistic score.
[1147] The emotional score is a numerical representation of the user's emotional response.
[1148] To fine-tune the artistic score, the emotional score is taken into account in the final evaluation.
[1149] Step 6: Delivering results
[1150] The terminal obtains the analysis results (art scores) in real time.
[1151] The terminal provides a display device for displaying the analysis results to the competitors and judges.
[1152] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[1153] The server records the analysis results and stores them in a database for future analysis and review.
[1154] In this way, each step works together to enable the system of the present invention to provide fair and consistent evaluations, and also to take into account the user's emotional responses.
[1155] Example 2
[1156] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1157] Conventional evaluation systems for artistic competitions and performances have difficulty providing fair and consistent evaluations. In particular, they lack systems that reflect the emotional reactions of audiences and judges in their evaluations, resulting in a lack of objectivity in the evaluations. Furthermore, they lack the ability to provide evaluations in real time and to store evaluation results.
[1158] The identification processing by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving video data or audio data, means for preprocessing the received video data or audio data, means for analyzing the preprocessed data and calculating an artistic score, means for recognizing the emotions of the contestants and spectators, means for fine-tuning the artistic score based on the emotion recognition results, and means for displaying the calculated artistic score. This makes it possible to provide fair and consistent evaluations and realize evaluations that take into account the emotional reactions of spectators and judges.
[1159] "Video data" refers to data that includes image information obtained from a camera, video input source, or the like.
[1160] "Audio data" refers to data containing sound information obtained from a microphone, audio input source, or the like.
[1161] "Preprocessing" refers to performing initial processing to extract specific features from received video and audio data.
[1162] "Analysis" refers to the evaluation of pre-processed data using AI models or other means to produce a specific result (e.g., artistic score).
[1163] "Artistic Score" is a score calculated based on the received video and audio data, or the emotional response of the competitors and spectators.
[1164] "Emotion recognition" is the process of analyzing the facial expressions, voice, body language, etc. of athletes and spectators to identify their emotional state.
[1165] "Fine-tuning based on emotion recognition results" means adjusting the aforementioned analysis results (artistic score) based on the data obtained through emotion recognition.
[1166] "Display" refers to the visual presentation of the analyzed results and fine-tuned artistic points.
[1167] The system of the present invention realizes a series of processes from receiving video or audio data, to preprocessing, analysis, display, and further fine-tuning evaluation by emotion recognition. This section explains the specific implementation method.
[1168] This system mainly consists of a server, a terminal, an emotion engine, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user. The emotion engine recognizes the user's emotions and fine-tunes the evaluation based on their emotional response.
[1169] Data collection
[1170] First, a user hosts an artistic competition or performance event (e.g., figure skating or piano competition). The server receives real-time video and audio data from the event. Video data comes from cameras and other video input sources, and audio data comes from microphones and other audio input sources.
[1171] Data Preprocessing
[1172] The server preprocesses the received video and audio data. For video data, motion capture and image recognition technology are used to extract the poses and movement characteristics of the person. For audio data, audio signal processing technology is used to extract musical features such as the pitch, intensity, and tempo of the sound. For example, the height of a figure skater's jump or the number of rotations in a spin, or the accuracy and rhythmic consistency of a piano performance can be obtained.
[1173] Data analysis
[1174] The server inputs the preprocessed data into an AI model to calculate artistic scores. In the case of video data, the AI model evaluates jump height, number of rotations, spin speed, etc., and in the case of audio data, it evaluates sound accuracy, expressiveness, rhythmic consistency, etc. The analysis results (artistic scores) obtained are stored in a database.
[1175] emotion recognition
[1176] The emotion engine analyzes the facial expressions, voice, and body language of competitors, performers, and audience members to identify their emotional state. The server uses the output of the emotion engine to fine-tune the artistic score, enabling a fair evaluation that reflects not only pure technical evaluation but also the emotional reactions of the audience and judges.
[1177] Providing results
[1178] The device acquires the analysis results (artistic scores) in real time and displays them to the competitors and judges. The user can check the displayed analysis results and use them to evaluate the competition and performance. The server also records the analysis results and stores them in a database for future analysis and review.
[1179] Specific examples
[1180] In the case of figure skating: A user hosts a competition, and the server receives live footage of the competition and collects data on jumps and spins. The server preprocesses this data, analyzes the jump height and spin rate using an AI model, and calculates and adjusts the artistic score. The results are displayed on the device in real time and provided to competitors and judges.
[1181] In the case of piano performance: A user hosts a piano competition, and the server receives the audio of the competition performance. The server preprocesses the audio data and analyzes pitch, intensity, tempo, etc. using an AI model to evaluate the accuracy of the sound and the technical accuracy of the performance. The emotion engine analyzes the facial expressions of the audience and judges and fine-tunes the artistic score based on their emotions. The results are displayed on the device in real time and provided to the user and judges.
[1182] Example prompt sentence:
[1183] Explain the processing steps of a system that evaluates the height of figure skating jumps and the number of rotations in spins, and adjusts the final artistic score taking into account the emotional response of the audience.
[1184]
[1185] The system analyzes musical characteristics such as pitch, intensity, and tempo during piano performances, and fine-tunes its evaluation based on the emotional responses of the audience and judges.
[1186] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1187] Step 1: Collect data
[1188] A user hosts events such as figure skating and piano competitions.
[1189] The server receives video or audio data from the event in real time. Specifically, video data is obtained from a camera or video input, and audio data is obtained from a microphone or audio input.
[1190] Input: Video data from a camera or video input, audio data from a microphone or audio input
[1191] Output: Raw data stored on the server
[1192] Step 2: Data Preprocessing
[1193] The server preprocesses the video and audio data received.
[1194] In the case of video data, motion capture and image recognition technology are used to extract a person's pose and movement characteristics.
[1195] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[1196] Input: Received raw video and audio data
[1197] Output: Preprocessed feature data (e.g., person pose data, sound characteristic data)
[1198] Step 3: Data analysis
[1199] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[1200] When using video data, the AI model evaluates factors such as jump height, number of rotations, and spin speed.
[1201] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[1202] Input: Preprocessed feature data
[1203] Output: Art score calculated by the AI model
[1204] Step 4: Emotion Recognition
[1205] The emotion engine analyzes the facial expressions, voice or body language of athletes, performers and spectators to identify their emotional state.
[1206] The server fine-tunes the artistic score based on the output of the emotion engine.
[1207] Input: Emotional data obtained from the camera or microphone (facial expressions, voice, body language)
[1208] Output: Emotional state analyzed by emotion engine, fine-tuned artistic score
[1209] Step 5: Delivering results
[1210] The terminal obtains the analysis results (art scores) in real time.
[1211] The terminal displays the analysis results to the competitors and judges.
[1212] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[1213] The server records the analysis results and stores them in a database for future analysis and review.
[1214] Input: Fine-tuned art scores, emotion recognition results
[1215] Output: Final evaluation results displayed in real time on the device, evaluation results stored in the database
[1216] (Application example 2)
[1217] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1218] Currently, there are almost no systems in physical stores that can grasp customer emotions and behavior in real time and provide appropriate responses based on that information. This makes it difficult to provide services that immediately reflect customer needs and emotions, limiting the improvement of customer satisfaction. It is also difficult to introduce a fair and consistent evaluation system.
[1219] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving video data or audio data, means for preprocessing the received video data or audio data, means for analyzing the preprocessed data and calculating an artistic score, means for displaying the calculated artistic score, and means for recognizing the user's emotions and fine-tuning the evaluation based on the emotional response. This makes it possible to analyze customer emotions and behavior in real time in physical stores and provide appropriate feedback to store staff immediately. Furthermore, using an emotion recognition engine enables a fair and consistent evaluation system and improves customer satisfaction.
[1220] "Video data" refers to image or video data acquired through a camera, video device, or the like.
[1221] "Audio data" refers to sound or voice data acquired through a microphone, recording device, or the like.
[1222] "Receiving means" refers to a function or system that allows a server or terminal to receive video data or audio data from the outside.
[1223] "Preprocessing means" refers to a function or system that processes received video data and audio data and converts them into a format suitable for analysis.
[1224] "Analysis means" refers to a function or system for calculating evaluation points based on preprocessed data.
[1225] "Artistic score" is a numerical evaluation of video and audio, and is an index used to evaluate technical and artistic aspects.
[1226] "Display means" refers to a function or system for visually displaying the calculated evaluation scores and analysis results.
[1227] "Emotion recognition" is a technology that identifies a user's emotional state based on data such as facial expressions and voice.
[1228] "Fine-tuning of the evaluation" is the process of adjusting the evaluation score calculated based on the emotion recognition results.
[1229] A "camera" is a photographing device for acquiring video data.
[1230] A "microphone" is a recording device for capturing audio data.
[1231] An "emotion engine" is software or hardware for analyzing a user's emotional state.
[1232] A "physical store" is a sales location set up in a physical location where customers visit in person to make purchases.
[1233] A "store clerk" is an employee who provides services to customers in a physical store.
[1234] The system of the present invention recognizes customer emotions and provides feedback to store clerks in real time. This system consists of a server, a camera, a microphone, a terminal, and an emotion engine. Specific hardware and software used include OpenCV, Keras, librosa, etc.
[1235] The system is configured as follows:
[1236] 1. Data Collection:
[1237] The server receives video and audio data from cameras and microphones installed in the store. The cameras capture customers' facial expressions in real time, and the microphones record their voices.
[1238] 2. Data preprocessing:
[1239] The received video and audio data are preprocessed by the server. The video data undergoes face recognition and facial expression analysis using OpenCV. The audio data undergoes audio feature extraction (e.g., MFCC) using librosa.
[1240] 3. Data Analysis:
[1241] The pre-processed data is fed into a generative AI model using Keras to analyze the customer's emotional state. Emotions are recognized from facial and vocal characteristics, and specific emotional responses (e.g., joy, anger, sadness, etc.) are identified by an emotion engine.
[1242] 4. Providing results:
[1243] The server displays the analysis results on the terminal in real time. The store clerk can then use this information to respond appropriately to the customer. For example, if the customer is angry, a more careful response appropriate to the situation will be recommended. If the customer is satisfied, the usual service will be provided.
[1244] As a concrete example, we will show an example of inputting a prompt sentence into a generative AI model.
[1245] Prompt: "Implement a system in a brick-and-mortar store that recognizes customer emotions and instructs store associates on appropriate responses based on facial and voice data. Key scenarios to consider include when a customer is angry or sad."
[1246] The above configuration and procedures make it possible to improve the customer experience in physical stores, and are expected to increase customer satisfaction.
[1247] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1248] Step 1: Collect data
[1249] The server receives video and audio data of customers in the store via a camera and microphone. Specifically, the camera captures the customer's face and body movements, and the microphone records the customer's voice. The input is video data from the camera and audio data from the microphone, which are then sent to the server.
[1250] Step 2: Preprocessing the data
[1251] The server preprocesses the received video and audio data. Specifically, it uses OpenCV to recognize faces and facial expressions from the video data and normalizes the images. For audio data, it uses librosa to convert the audio signal into features (e.g., MFCCs). The input is raw video and audio data, and the output is preprocessed face image data and audio feature data.
[1252] Step 3: Analyze the data
[1253] The server analyzes the preprocessed data. Specifically, a generative AI model using Keras is used to estimate the emotional state from the facial image data and the emotional state of the voice from the voice feature data. The input is the preprocessed facial image data and voice feature data, and the output is the customer's emotional state (e.g., joy, anger, sadness, etc.).
[1254] Step 4: Applying the Emotion Engine
[1255] The server further processes the analysis results using an emotion engine to refine the accuracy. The emotion engine refines the overall emotion rating based on the analyzed emotional state. The input is the analyzed emotional state, and the output is a fine-tuned emotion rating.
[1256] Step 5: View the results
[1257] The server displays the adjusted emotion evaluation on the terminal in real time. Specifically, it notifies the store clerk of the analysis results and the fine-tuned emotion evaluation. The input is the fine-tuned emotion evaluation, and the output is a notification message displayed on the store clerk's terminal. For example, if the customer is angry, the message displayed is "The customer is angry. Please be careful how you respond."
[1258] Through these steps, the system can recognize customer emotions in real time in physical stores and provide immediate and appropriate feedback to store staff.
[1259] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1260] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1261] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1262] [Fourth embodiment]
[1263] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1264] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1265] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1266] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1267] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1268] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1269] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1270] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1271] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1272] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1273] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1274] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1275] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1276] The system of the present invention analyzes video or audio data in real time to provide fair and consistent evaluation in artistic scoring competitions. The system has a series of functions for receiving, preprocessing, and analyzing video or audio data, and displaying the resulting artistic scores. The specific operation of the system is described below.
[1277] System Overview
[1278] This system consists of a server, a terminal, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user.
[1279] 1. Data collection
[1280] Users host events such as figure skating and piano competitions.
[1281] A server receives video data (e.g., video of a figure skating performance) and audio data (e.g., audio of a piano performance) from an event in real time.
[1282] System operation details
[1283] 1. Data Collection
[1284] A server receives a live stream of a figure skating or piano competition.
[1285] In the case of video data, the server acquires data in real time from cameras and other video input sources.
[1286] For audio data, the server captures data in real time from a microphone or other audio input source.
[1287] 2. Data Preprocessing
[1288] The server converts the received video and audio data into an easy-to-understand format.
[1289] In the case of figure skating, motion capture and image recognition technology are used to extract a person's poses and movement characteristics.
[1290] In the case of piano performance, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[1291] 3. Data Analysis
[1292] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[1293] In figure skating, the height of the jump, the number of rotations, and the speed of the spins are evaluated.
[1294] Piano performance is assessed on accuracy of tone, expressiveness, and rhythmic consistency.
[1295] The analyzed results (art scores) are stored in a database.
[1296] 4. Providing results
[1297] The terminal displays the analysis results (art scores) in real time.
[1298] The terminal is equipped with a display device for displaying information to competitors and judges.
[1299] The server records the results and makes them available for later analysis and review.
[1300] Specific examples
[1301] figure skating
[1302] 1. A user hosts a figure skating competition.
[1303] 2. The server receives live footage of the competition and collects data on jumps and spins.
[1304] 3. The server preprocesses the collected video data to extract information such as jump height and spin rotations.
[1305] 4. The server analyzes the preprocessed data using an AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[1306] 5. The device displays the analysis results in real time and provides them to the competitors and judges.
[1307] Piano performance
[1308] 1. A user hosts a piano competition.
[1309] 2. The server receives the audio of the competition performance.
[1310] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[1311] 4. The server analyzes the preprocessed data using an AI model to evaluate the sound characteristics and technical accuracy of the performance.
[1312] 5. The device displays the analysis results in real time and provides them to the user and judges.
[1313] In this way, the system of the present invention is able to provide fair and consistent grading for artistic competitions and reduce variability in grading.
[1314] The processing flow will be explained below.
[1315] Processing flow and each processing step
[1316] Step 1: Receiving Data
[1317] A user hosts a figure skating competition or a piano competition.
[1318] A server receives video or audio data from an event in real time.
[1319] In the case of video data, the data is obtained from a camera or video input source.
[1320] For audio data, data is obtained from a microphone or audio input source.
[1321] Step 2: Save your data
[1322] Create a directory for the server to temporarily store the data it receives.
[1323] The server temporarily saves the received data stream in the "raw_data" directory.
[1324] Step 3: Preprocessing the data
[1325] The server reads the saved data from the "raw_data" directory.
[1326] In the case of video data, motion capture and image recognition technology are used to extract people's poses and movement characteristics.
[1327] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[1328] The server saves the preprocessed data in the "processed_data" directory.
[1329] Step 4: Analyze the data
[1330] The server reads the preprocessed data from the "processed_data" directory.
[1331] The server inputs the preprocessed data into an AI model to calculate the artistic score.
[1332] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[1333] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[1334] The server stores the analysis results (art scores) in a database.
[1335] Step 5: Delivering results
[1336] The terminal obtains the analysis results (art scores) in real time.
[1337] The terminal provides a display device for displaying the analysis results to the competitors and judges.
[1338] The user can check the displayed analysis results and use them to evaluate the competition or performance.
[1339] The server records the analysis results and stores them in a database for future analysis and review.
[1340] In this way, with each step working together, the system of the present invention provides fair and consistent evaluation of artistic competitions and performances.
[1341] Example 1
[1342] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1343] Conventional evaluation methods for artistic competitions often rely on the subjective judgment and experience of judges, resulting in a lack of fairness and consistency. Real-time evaluation is also difficult, resulting in variations and delays in evaluation. The present invention aims to solve these problems and provide a system that provides fair and consistent evaluation in real time.
[1344] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1345] In this invention, the server includes a means for receiving video data or audio data, a means for preprocessing the received video data or audio data, and a means for inputting the preprocessed data into a generative AI model to calculate artistic scores, thereby enabling the server to analyze the video data or audio data in real time and provide fair and consistent evaluations.
[1346] "Video Data" refers to visual information of an action or scene captured from a camera or other video input device.
[1347] "Audio Data" refers to sound information collected from a microphone or other audio input device.
[1348] "Means for receiving" refers to the technical means by which the server retrieves video and audio data from a network or device.
[1349] "Pre-processing means" refers to technical means for converting received video and audio data into an understandable format. Examples include motion capture and audio signal processing techniques.
[1350] A "generative AI model" is a model that uses artificial intelligence technology to analyze data and generate scores or ratings for specific purposes.
[1351] "Means for calculating artistic scores" refers to the technical means for inputting preprocessed data into a generative AI model and conducting a quantitative evaluation.
[1352] "Means for displaying" refers to the technical means for visually presenting the artistic score, which is the result of the analysis, to users and judges.
[1353] "Personal action" refers to a figure skating routine or other artistic movement performed by a particular person.
[1354] "Extracting pose and movement features" refers to the act of analyzing specific postures and movements from video data and extracting related features.
[1355] "Musical performance" refers to the act of performing a musical piece using instruments and voices.
[1356] "Extracting musical features" refers to the act of analyzing attributes such as pitch, intensity, and tempo of sound from audio data and extracting features.
[1357] The system of the present invention uses video and audio data analysis technology to provide fair and consistent artistic evaluation in real time. The system is mainly composed of three elements: a server, a terminal, and a user.
[1358] server
[1359] Data reception: The server first receives video and audio data from events such as figure skating and piano competitions. Specifically, it acquires streaming data via cameras and microphones using RTSP (Real-Time Streaming Protocol) or similar.
[1360] Data preprocessing: The server then preprocesses the received data. For video data, motion capture and image recognition technologies are used to extract human poses and movement characteristics. This is done using libraries such as OpenPose. For audio data, FFT (Fast Fourier Transform) is used to extract musical features such as pitch, intensity, and tempo.
[1361] Data Analysis: The preprocessed data is then fed into a generative AI model. The server uses deep learning models such as TensorFlow or PyTorch to analyze this data and calculate artistic scores. In the case of figure skating, scores are generated based on factors such as jump height, number of rotations, and spin speed. In the case of piano playing, scores are evaluated based on pitch accuracy, tempo consistency, and expressiveness.
[1362] Result storage: The analysis results (art scores) are stored in a database by the server, allowing for later analysis and review. This can be done using a SQL or NoSQL database.
[1363] Terminal
[1364] Display of results: The device displays the analysis results obtained from the server in real time. An interface is provided to visually present the results to competitors and judges. Specifically, the results are displayed using a web browser or a mobile app.
[1365] Accepting user input: The terminal also provides an interface for accepting user input, allowing operations such as starting and stopping a competition and checking results.
[1366] User
[1367] Hosting an event: A user hosts an event such as a figure skating or piano competition. They distribute entry forms to recruit participants, schedule the competition, and notify them.
[1368] Specific examples
[1369] In the case of figure skating
[1370] 1. A user hosts a figure skating competition.
[1371] Example: A user posts a competition and gathers participants online.
[1372] 2. The server receives live footage of the competition and collects data on jumps and spins.
[1373] Specific operation: The server receives video streaming data from the camera and extracts motion features using OpenPose.
[1374] 3. The server analyzes the preprocessed data using a generative AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[1375] Example prompt: "Analyze video of a figure skating competition and measure the jump height, number of rotations, and spin speed. Calculate an artistic score based on this data."
[1376] 4. The device displays the analysis results in real time and provides them to the competitors and judges.
[1377] Example: Results are displayed instantly on large screens at the stadium.
[1378] For piano performances
[1379] 1. A user hosts a piano competition.
[1380] Specific actions: Recruit participants using an entry form.
[1381] 2. The server receives the audio of the competition performance.
[1382] Example: A server captures audio data from a microphone in real time.
[1383] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[1384] Example: Extracting sound characteristics using FFT.
[1385] 4. The server analyzes the preprocessed data using a generative AI model to evaluate the sound characteristics and technical accuracy of the performance.
[1386] 5. The device displays the analysis results in real time and provides them to the user and judges.
[1387] Example: Analysis results are displayed instantly on smartphone apps and web apps.
[1388] In this way, the system of the present invention is able to provide fair and consistent grading for artistic competitions and reduce variability in grading.
[1389] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1390] Step 1: Collect data
[1391] A user hosts an event such as a figure skating or piano competition. Specifically, the user prepares an entry form, recruits participants, and sets and notifies the schedule of the competition. This completes the preparations for the competition.
[1392] The server receives video and audio data from the event in real time. Specific inputs include streaming video and audio data captured using a camera or microphone. The raw data is then stored on the server.
[1393] Step 2: Preprocessing the data
[1394] The server preprocesses the video and audio data it receives. Specifically, the server receives raw data and converts it into a format that is easy to analyze.
[1395] For video data, the server uses motion capture and image recognition technology to extract poses and movement characteristics of people, and specific outputs, such as jump height and spin rotations, are extracted using libraries such as OpenPose.
[1396] For audio data, the server uses FFT (Fast Fourier Transform) to extract musical features such as pitch, intensity, and tempo. Specific outputs include pitch accuracy, intensity distribution, and tempo fluctuation data.
[1397] Step 3: Analyze the data
[1398] The server inputs the preprocessed data into a generative AI model to calculate the artistic score. Specific inputs include preprocessed motion data and musical features. This is then input into an AI model such as TensorFlow or PyTorch.
[1399] In the case of figure skating, the server evaluates the height of jumps, number of rotations, speed of spins, etc. The output is a score for each element and an overall artistic score.
[1400] For piano performances, the server evaluates pitch accuracy, tempo consistency, and expressiveness, producing an output that measures the overall performance's technical accuracy and artistic evaluation score.
[1401] Step 4: Save the results
[1402] The server stores the analysis results (art scores) in a database. Specific inputs include the generated art scores and scores for each evaluation item. These are recorded in an SQL or NoSQL database.
[1403] As an output, the analysis results are stored in a format that can be used for later analysis and review, for example to analyze historical and trending results.
[1404] Step 5: View the results
[1405] The device displays the analysis results (art scores) in real time. Specific inputs include analysis results obtained from the server, which are then displayed on the user interface of a web browser or mobile app.
[1406] The output allows athletes and judges to check the evaluation results in real time. For example, the results may be displayed on a large screen at the competition venue or instantly on a smartphone app.
[1407] Step 6: User feedback
[1408] The user checks the evaluation results through the terminal and provides feedback. Specific inputs include feedback data from the user, which is sent to the server and stored.
[1409] As an output, the feedback is used to refine the system and adjust the evaluation criteria for the next iteration, thereby improving the overall accuracy of the system and the user experience.
[1410] (Application example 1)
[1411] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1412] With conventional methods, evaluation of live performances in physical venues is often subjective, making it difficult to achieve fair and consistent evaluations. Furthermore, the delay in real-time evaluation feedback often makes performers and audiences feel that the evaluations lack transparency and immediacy.
[1413] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1414] In this invention, the server includes means for receiving video data or audio data, means for pre-processing the received video data or audio data, means for analyzing the pre-processed data and calculating an artistic score, means for displaying the calculated artistic score, and means for capturing video and audio data of live performances through cameras and microphones in the store, thereby enabling fair and consistent real-time evaluation of live performances in physical stores.
[1415] "Video data" refers to dynamic or static image information acquired through a visual device such as a camera.
[1416] "Audio data" refers to an audio signal obtained through an audio device such as a microphone.
[1417] "Preprocessing" refers to a series of data manipulation steps that convert the received raw data into a format that is easier to analyze.
[1418] "Art score calculation" refers to the process of analyzing pre-processed data and generating a score based on specific evaluation criteria.
[1419] "Display means" refers to a display device or software interface that visually presents the analysis results to the user.
[1420] "In-store cameras" refer to video capture devices installed in physical stores to capture video data of live performances.
[1421] "In-store microphones" refer to audio capture devices installed in physical stores to capture audio data from live performances.
[1422] The system embodying this invention evaluates live performance in a physical store in real time and provides fair and consistent scoring. The system mainly uses a server, in-store cameras and microphones, a smartphone app, and a display.
[1423] Hardware and Software Configuration
[1424] 1. Server:
[1425] Role: Receiving, preprocessing, analyzing and storing data.
[1426] Software used:
[1427] Motion capture technology: OpenPose, etc.
[1428] Audio signal processing technology: librosa, etc.
[1429] AI analysis model: TensorFlow or PyTorch
[1430] 2. In-store cameras:
[1431] Role: Capture video data of the performance.
[1432] Example: High-definition video camera
[1433] 3. In-store microphones:
[1434] Role: Captures audio data of the performance.
[1435] Example: High-sensitivity microphone
[1436] 4. Smartphone App:
[1437] Role: Displays the score and accepts user input.
[1438] Example: Android or iOS app
[1439] 5. Display:
[1440] Role: Displaying evaluation results in-store.
[1441] Example: Large displays in stores
[1442] A natural language description of the program's operation
[1443] The server captures video data of the performance in real time through in-store cameras, and simultaneously captures audio data through in-store microphones. Motion capture technology is used to extract the poses and movement characteristics of the performers from the received video data, and audio signal processing technology is used to extract musical features such as pitch, intensity, and tempo from the audio data. The extracted features are input into an AI analysis model to calculate the performance's artistic score. The calculated artistic score is saved on the server and displayed in real time on a smartphone app and on in-store displays.
[1444] Specific examples
[1445] For example, imagine a live music concert is being held in a store. When the live performance begins, cameras in the store collect video data and microphones collect audio data, which are then sent to a server. The server preprocesses the received video data and analyzes the performers' movement characteristics using OpenPose. Musical features are extracted from the audio data using librosa, and these are input into an AI model built with TensorFlow and PyTorch. The model generates a score based on evaluation criteria, and the score is displayed in real time on a smartphone app or display.
[1446] Prompt Sentence Examples
[1447] Please use an AI model to analyze the beauty of jumps and spins in figure skating performance footage and provide a real-time evaluation score. Please enter the video data below. Please display the analysis results.
[1448] <Video data>
[1449] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1450] Step 1:
[1451] Data collection
[1452] When a live event begins, cameras and microphones in the store capture video and audio data of the performance in real time. These data are sent to a server. The cameras provide video data (e.g., the performers' movements) and the microphones provide audio data (e.g., the musical performance).
[1453] Input: Live performance video and audio data
[1454] Output: Video and audio data sent to the server
[1455] Step 2:
[1456] Data Preprocessing
[1457] The server preprocesses the received video data using OpenPose to extract human poses and movement features, and preprocesses the audio data using librosa to extract musical features such as pitch, intensity, and tempo, resulting in an analyzable dataset.
[1458] Input: Video and audio data sent to the server
[1459] Output: Preprocessed motion feature data and musical feature data
[1460] Step 3:
[1461] Data analysis
[1462] The server inputs the preprocessed data into an AI analysis model (e.g., using TensorFlow or PyTorch) to calculate the artistic score of the performance. The analysis model uses features such as the accuracy of the movements, the rhythm of the music, and the tempo as evaluation criteria.
[1463] Input: Preprocessed motion feature data and musical feature data
[1464] Output: Calculated art score
[1465] Step 4:
[1466] Saving the results
[1467] The server stores the calculated artistic scores in a database, which can be used for later review and analysis.
[1468] Input: Calculated art score
[1469] Output: Art points stored in the database
[1470] Step 5:
[1471] Displaying the results
[1472] The device displays the analysis results in real time, i.e., artistic scores, on a smartphone app and on displays in the store, allowing performers and audience members to immediately check the evaluation results.
[1473] Input: Art points stored in the database
[1474] Output: Smartphone app and artwork displayed on the screen
[1475] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1476] The system of the present invention includes a series of processes from receiving video or audio data to pre-processing, analysis, and display, and is combined with an emotion engine that recognizes the user's emotions. This system not only enables fair and consistent evaluation in artistic competitions and performances, but also takes into account the user's emotional responses.
[1477] System Overview
[1478] This system consists of a server, a terminal, an emotion engine, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user. The emotion engine recognizes the user's emotions and fine-tunes the evaluation based on their emotional response.
[1479] 1. Data collection
[1480] Users host figure skating and piano competitions.
[1481] The server receives video data (e.g., video of a figure skating performance) and audio data (e.g., audio of a piano performance) from the event in real time.
[1482] System operation details
[1483] 1. Data Collection
[1484] The server receives the video and audio data in real time.
[1485] In the case of video data, the data is obtained from a camera or video input source.
[1486] For audio data, data is obtained from a microphone or audio input source.
[1487] 2. Data Preprocessing
[1488] The server preprocesses the received video and audio data.
[1489] In the case of video data, motion capture and image recognition technology are used to extract a person's pose and movement characteristics.
[1490] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[1491] 3. Data Analysis
[1492] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[1493] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[1494] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[1495] The analysis results (art scores) obtained are stored in a database.
[1496] 4. Emotion recognition
[1497] An emotion engine analyzes the facial expressions, vocal or body language of athletes, performers and spectators to identify their emotional state.
[1498] The server uses the output of the emotion engine to fine-tune the artistic score.
[1499] 5. Providing results
[1500] The terminal obtains the analysis results (art scores) in real time.
[1501] The terminal displays the analysis results to the competitor and judges.
[1502] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[1503] The server records the analysis results and stores them in a database for future analysis and review.
[1504] Specific examples
[1505] figure skating
[1506] 1. A user hosts a figure skating competition.
[1507] 2. The server receives live footage of the competition and collects data on jumps and spins.
[1508] 3. The server preprocesses the collected video data to extract information such as jump height and spin rotations.
[1509] 4. The server analyzes the preprocessed data using an AI model to calculate the success rate of the jump and artistic score based on the beauty of the jump.
[1510] 5. The emotion engine analyzes the facial expressions of the audience and judges and measures the level of emotion based on each element.
[1511] 6. The server fine-tunes the artistic score based on the output of the emotion engine.
[1512] 7. The device displays the analysis results in real time and provides them to the competitors and judges.
[1513] Piano performance
[1514] 1. A user hosts a piano competition.
[1515] 2. The server receives the audio of the competition performance.
[1516] 3. The server preprocesses the stored audio data and extracts musical features such as pitch, intensity, and tempo.
[1517] 4. The server analyzes the preprocessed data using an AI model to evaluate the sound characteristics and technical accuracy of the performance.
[1518] 5. The emotion engine analyzes the facial expressions of the audience and judges and measures the level of emotion based on each element.
[1519] 6. The server fine-tunes the artistic score based on the output of the emotion engine.
[1520] 7. The device displays the analysis results in real time and provides them to the user and judges.
[1521] In this way, each step works together to enable the system of the present invention to provide fair and consistent evaluations, and also to take into account the user's emotional responses.
[1522] The processing flow will be explained below.
[1523] Processing flow and each processing step
[1524] Step 1: Receiving Data
[1525] A user hosts a figure skating competition or a piano competition.
[1526] A server receives video or audio data from an event in real time.
[1527] In the case of video data, the data is obtained from a camera or video input source.
[1528] For audio data, data is obtained from a microphone or audio input source.
[1529] Step 2: Save your data
[1530] Create a directory for the server to temporarily store the data it receives.
[1531] The server temporarily saves the received data stream in the "raw_data" directory.
[1532] Step 3: Preprocessing the data
[1533] The server reads the saved data from the "raw_data" directory and performs preprocessing.
[1534] In the case of video data, motion capture and image recognition technology are used to extract people's poses and movement characteristics.
[1535] Motion capture is a technology that identifies the joint positions of a person in a video.
[1536] Image recognition technology is a technology that recognizes specific movements (e.g., jumps and spins).
[1537] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[1538] Pitch is a feature that indicates the pitch of a sound.
[1539] Intensity is a feature that indicates the loudness of a sound.
[1540] Tempo is a characteristic that indicates the speed of a performance.
[1541] The server saves the preprocessed data in the "processed_data" directory.
[1542] Step 4: Analyze the data
[1543] The server reads the preprocessed data from the "processed_data" directory.
[1544] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[1545] In the case of video data, the height of the jump, number of rotations, spin speed, etc. are evaluated.
[1546] The jump height refers to the height from the start point of the jump to the landing point.
[1547] The number of rotations refers to the number of rotations during the jump.
[1548] The rotation speed of the spin indicates the speed of rotation during the spin.
[1549] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[1550] The accuracy of the sound is evaluated by whether the sound is produced according to the musical score.
[1551] Expressiveness is measured by the performer's richness of emotion and expression.
[1552] Rhythmic consistency is measured by whether the performance maintains a consistent tempo.
[1553] The server stores the analysis results (art scores) in a database.
[1554] Step 5: Emotion Recognition
[1555] The emotion engine analyzes the facial expressions, voice or body language of athletes and spectators to identify their emotional state.
[1556] Facial expression analysis is a technology that analyzes facial expressions from camera footage and identifies emotions.
[1557] Voice analysis is a technology that analyzes the tone and strength of a voice to identify emotions.
[1558] Body language analysis is a technique that analyzes gestures and postures to identify emotions.
[1559] The server uses the output of the emotion engine to fine-tune the artistic score.
[1560] The emotional score is a numerical representation of the user's emotional response.
[1561] To fine-tune the artistic score, the emotional score is taken into account in the final evaluation.
[1562] Step 6: Delivering results
[1563] The terminal obtains the analysis results (art scores) in real time.
[1564] The terminal provides a display device for displaying the analysis results to the competitors and judges.
[1565] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[1566] The server records the analysis results and stores them in a database for future analysis and review.
[1567] In this way, each step works together to enable the system of the present invention to provide fair and consistent evaluations, and also to take into account the user's emotional responses.
[1568] Example 2
[1569] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1570] Conventional evaluation systems for artistic competitions and performances have difficulty providing fair and consistent evaluations. In particular, they lack systems that reflect the emotional reactions of audiences and judges in their evaluations, resulting in a lack of objectivity in the evaluations. Furthermore, they lack the ability to provide evaluations in real time and to store evaluation results.
[1571] The identification processing by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving video data or audio data, means for preprocessing the received video data or audio data, means for analyzing the preprocessed data and calculating an artistic score, means for recognizing the emotions of the contestants and spectators, means for fine-tuning the artistic score based on the emotion recognition results, and means for displaying the calculated artistic score. This makes it possible to provide fair and consistent evaluations and realize evaluations that take into account the emotional reactions of spectators and judges.
[1572] "Video data" refers to data that includes image information obtained from a camera, video input source, or the like.
[1573] "Audio data" refers to data containing sound information obtained from a microphone, audio input source, or the like.
[1574] "Preprocessing" refers to performing initial processing to extract specific features from received video and audio data.
[1575] "Analysis" refers to the evaluation of pre-processed data using AI models or other means to produce a specific result (e.g., artistic score).
[1576] "Artistic Score" is a score calculated based on the received video and audio data, or the emotional response of the competitors and spectators.
[1577] "Emotion recognition" is the process of analyzing the facial expressions, voice, body language, etc. of athletes and spectators to identify their emotional state.
[1578] "Fine-tuning based on emotion recognition results" means adjusting the aforementioned analysis results (artistic score) based on the data obtained through emotion recognition.
[1579] "Display" refers to the visual presentation of the analyzed results and fine-tuned artistic points.
[1580] The system of the present invention realizes a series of processes from receiving video or audio data, to preprocessing, analysis, display, and further fine-tuning evaluation by emotion recognition. This section explains the specific implementation method.
[1581] This system mainly consists of a server, a terminal, an emotion engine, and a user. The server receives, preprocesses, analyzes, and stores data, while the terminal displays the results and accepts input from the user. The emotion engine recognizes the user's emotions and fine-tunes the evaluation based on their emotional response.
[1582] Data collection
[1583] First, a user hosts an artistic competition or performance event (e.g., figure skating or piano competition). The server receives real-time video and audio data from the event. Video data comes from cameras and other video input sources, and audio data comes from microphones and other audio input sources.
[1584] Data Preprocessing
[1585] The server preprocesses the received video and audio data. For video data, motion capture and image recognition technology are used to extract the poses and movement characteristics of the person. For audio data, audio signal processing technology is used to extract musical features such as the pitch, intensity, and tempo of the sound. For example, the height of a figure skater's jump or the number of rotations in a spin, or the accuracy and rhythmic consistency of a piano performance can be obtained.
[1586] Data analysis
[1587] The server inputs the preprocessed data into an AI model to calculate artistic scores. In the case of video data, the AI model evaluates jump height, number of rotations, spin speed, etc., and in the case of audio data, it evaluates sound accuracy, expressiveness, rhythmic consistency, etc. The analysis results (artistic scores) obtained are stored in a database.
[1588] emotion recognition
[1589] The emotion engine analyzes the facial expressions, voice, and body language of competitors, performers, and audience members to identify their emotional state. The server uses the output of the emotion engine to fine-tune the artistic score, enabling a fair evaluation that reflects not only pure technical evaluation but also the emotional reactions of the audience and judges.
[1590] Providing results
[1591] The device acquires the analysis results (artistic scores) in real time and displays them to the competitors and judges. The user can check the displayed analysis results and use them to evaluate the competition and performance. The server also records the analysis results and stores them in a database for future analysis and review.
[1592] Specific examples
[1593] In the case of figure skating: A user hosts a competition, and the server receives live footage of the competition and collects data on jumps and spins. The server preprocesses this data, analyzes the jump height and spin rate using an AI model, and calculates and adjusts the artistic score. The results are displayed on the device in real time and provided to competitors and judges.
[1594] In the case of piano performance: A user hosts a piano competition, and the server receives the audio of the competition performance. The server preprocesses the audio data and analyzes pitch, intensity, tempo, etc. using an AI model to evaluate the accuracy of the sound and the technical accuracy of the performance. The emotion engine analyzes the facial expressions of the audience and judges and fine-tunes the artistic score based on their emotions. The results are displayed on the device in real time and provided to the user and judges.
[1595] Example prompt sentence:
[1596] Explain the processing steps of a system that evaluates the height of figure skating jumps and the number of rotations in spins, and adjusts the final artistic score taking into account the emotional response of the audience.
[1597]
[1598] The system analyzes musical characteristics such as pitch, intensity, and tempo during piano performances, and fine-tunes its evaluation based on the emotional responses of the audience and judges.
[1599] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1600] Step 1: Collect data
[1601] A user hosts events such as figure skating and piano competitions.
[1602] The server receives video or audio data from the event in real time. Specifically, video data is obtained from a camera or video input, and audio data is obtained from a microphone or audio input.
[1603] Input: Video data from a camera or video input, audio data from a microphone or audio input
[1604] Output: Raw data stored on the server
[1605] Step 2: Data Preprocessing
[1606] The server preprocesses the video and audio data received.
[1607] In the case of video data, motion capture and image recognition technology are used to extract a person's pose and movement characteristics.
[1608] In the case of audio data, audio signal processing techniques are used to extract musical features such as pitch, intensity, and tempo of the sound.
[1609] Input: Received raw video and audio data
[1610] Output: Preprocessed feature data (e.g., person pose data, sound characteristic data)
[1611] Step 3: Data analysis
[1612] The server inputs the preprocessed data into the AI model and calculates the artistic score.
[1613] When using video data, the AI model evaluates factors such as jump height, number of rotations, and spin speed.
[1614] In the case of audio data, the accuracy of the sound, expressiveness, rhythmic consistency, etc. are evaluated.
[1615] Input: Preprocessed feature data
[1616] Output: Art score calculated by the AI model
[1617] Step 4: Emotion Recognition
[1618] The emotion engine analyzes the facial expressions, voice or body language of athletes, performers and spectators to identify their emotional state.
[1619] The server fine-tunes the artistic score based on the output of the emotion engine.
[1620] Input: Emotional data obtained from the camera or microphone (facial expressions, voice, body language)
[1621] Output: Emotional state analyzed by emotion engine, fine-tuned artistic score
[1622] Step 5: Delivering results
[1623] The terminal obtains the analysis results (art scores) in real time.
[1624] The terminal displays the analysis results to the competitors and judges.
[1625] The user can check the displayed analysis results and use them to evaluate competitions and performances.
[1626] The server records the analysis results and stores them in a database for future analysis and review.
[1627] Input: Fine-tuned art scores, emotion recognition results
[1628] Output: Final evaluation results displayed in real time on the device, evaluation results stored in the database
[1629] (Application example 2)
[1630] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1631] Currently, there are almost no systems in physical stores that can grasp customer emotions and behavior in real time and provide appropriate responses based on that information. This makes it difficult to provide services that immediately reflect customer needs and emotions, limiting the improvement of customer satisfaction. It is also difficult to introduce a fair and consistent evaluation system.
[1632] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving video data or audio data, means for preprocessing the received video data or audio data, means for analyzing the preprocessed data and calculating an artistic score, means for displaying the calculated artistic score, and means for recognizing the user's emotions and fine-tuning the evaluation based on the emotional response. This makes it possible to analyze customer emotions and behavior in real time in physical stores and provide appropriate feedback to store staff immediately. Furthermore, using an emotion recognition engine enables a fair and consistent evaluation system and improves customer satisfaction.
[1633] "Video data" refers to image or video data acquired through a camera, video device, or the like.
[1634] "Audio data" refers to sound or voice data acquired through a microphone, recording device, or the like.
[1635] "Receiving means" refers to a function or system that allows a server or terminal to receive video data or audio data from the outside.
[1636] "Preprocessing means" refers to a function or system that processes received video data and audio data and converts them into a format suitable for analysis.
[1637] "Analysis means" refers to a function or system for calculating evaluation points based on preprocessed data.
[1638] "Artistic score" is a numerical evaluation of video and audio, and is an index used to evaluate technical and artistic aspects.
[1639] "Display means" refers to a function or system for visually displaying the calculated evaluation scores and analysis results.
[1640] "Emotion recognition" is a technology that identifies a user's emotional state based on data such as facial expressions and voice.
[1641] "Fine-tuning of the evaluation" is the process of adjusting the evaluation score calculated based on the emotion recognition results.
[1642] A "camera" is a photographing device for acquiring video data.
[1643] A "microphone" is a recording device for capturing audio data.
[1644] An "emotion engine" is software or hardware for analyzing a user's emotional state.
[1645] A "physical store" is a sales location set up in a physical location where customers visit in person to make purchases.
[1646] A "store clerk" is an employee who provides services to customers in a physical store.
[1647] The system of the present invention recognizes customer emotions and provides feedback to store clerks in real time. This system consists of a server, a camera, a microphone, a terminal, and an emotion engine. Specific hardware and software used include OpenCV, Keras, librosa, etc.
[1648] The system is configured as follows:
[1649] 1. Data Collection:
[1650] The server receives video and audio data from cameras and microphones installed in the store. The cameras capture customers' facial expressions in real time, and the microphones record their voices.
[1651] 2. Data preprocessing:
[1652] The received video and audio data are preprocessed by the server. The video data undergoes face recognition and facial expression analysis using OpenCV. The audio data undergoes audio feature extraction (e.g., MFCC) using librosa.
[1653] 3. Data Analysis:
[1654] The pre-processed data is fed into a generative AI model using Keras to analyze the customer's emotional state. Emotions are recognized from facial and vocal characteristics, and specific emotional responses (e.g., joy, anger, sadness, etc.) are identified by an emotion engine.
[1655] 4. Providing results:
[1656] The server displays the analysis results on the terminal in real time. The store clerk can then use this information to respond appropriately to the customer. For example, if the customer is angry, a more careful response appropriate to the situation will be recommended. If the customer is satisfied, the usual service will be provided.
[1657] As a concrete example, we will show an example of inputting a prompt sentence into a generative AI model.
[1658] Prompt: "Implement a system in a brick-and-mortar store that recognizes customer emotions and instructs store associates on appropriate responses based on facial and voice data. Key scenarios to consider include when a customer is angry or sad."
[1659] The above configuration and procedures make it possible to improve the customer experience in physical stores, and are expected to increase customer satisfaction.
[1660] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1661] Step 1: Collect data
[1662] The server receives video and audio data of customers in the store via a camera and microphone. Specifically, the camera captures the customer's face and body movements, and the microphone records the customer's voice. The input is video data from the camera and audio data from the microphone, which are then sent to the server.
[1663] Step 2: Preprocessing the data
[1664] The server preprocesses the received video and audio data. Specifically, it uses OpenCV to recognize faces and facial expressions from the video data and normalizes the images. For audio data, it uses librosa to convert the audio signal into features (e.g., MFCCs). The input is raw video and audio data, and the output is preprocessed face image data and audio feature data.
[1665] Step 3: Analyze the data
[1666] The server analyzes the preprocessed data. Specifically, a generative AI model using Keras is used to estimate the emotional state from the facial image data and the emotional state of the voice from the voice feature data. The input is the preprocessed facial image data and voice feature data, and the output is the customer's emotional state (e.g., joy, anger, sadness, etc.).
[1667] Step 4: Applying the Emotion Engine
[1668] The server further processes the analysis results using an emotion engine to refine the accuracy. The emotion engine refines the overall emotion rating based on the analyzed emotional state. The input is the analyzed emotional state, and the output is a fine-tuned emotion rating.
[1669] Step 5: View the results
[1670] The server displays the adjusted emotion evaluation on the terminal in real time. Specifically, it notifies the store clerk of the analysis results and the fine-tuned emotion evaluation. The input is the fine-tuned emotion evaluation, and the output is a notification message displayed on the store clerk's terminal. For example, if the customer is angry, the message displayed is "The customer is angry. Please be careful how you respond."
[1671] Through these steps, the system can recognize customer emotions in real time in physical stores and provide immediate and appropriate feedback to store staff.
[1672] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1673] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1674] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1675] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1676] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1677] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1678] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1679] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1680] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1681] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1682] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1683] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1684] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1685] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1686] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1687] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1688] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1689] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1690] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1691] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1692] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1693] The following is further disclosed regarding the above embodiment.
[1694] (Claim 1)
[1695] means for receiving video data or audio data;
[1696] means for pre-processing the received video or audio data;
[1697] means for analyzing the preprocessed data and calculating an art score;
[1698] a means for displaying the calculated artistic score;
[1699] A system including:
[1700] (Claim 2)
[1701] 2. The system according to claim 1, wherein the video data includes human movements, and the pre-processing means extracts poses and movement features of the human.
[1702] (Claim 3)
[1703] 2. The system of claim 1, wherein the audio data includes a musical performance, and the preprocessing means extracts musical features.
[1704] "Example 1"
[1705] (Claim 1)
[1706] means for receiving video data or audio data;
[1707] means for pre-processing the received video or audio data;
[1708] A means for inputting pre-processed data into a generative AI model to calculate artistic scores;
[1709] a means for displaying the calculated artistic score;
[1710] A system including:
[1711] (Claim 2)
[1712] 2. The system according to claim 1, wherein the video data includes human movements, and the pre-processing means extracts poses and movement features of the human.
[1713] (Claim 3)
[1714] 2. The system of claim 1, wherein the audio data includes a musical performance, and the preprocessing means extracts musical features.
[1715] "Application Example 1"
[1716] (Claim 1)
[1717] means for receiving video data or audio data;
[1718] means for pre-processing the received video or audio data;
[1719] means for analyzing the preprocessed data and calculating an art score;
[1720] a means for displaying the calculated artistic score;
[1721] A means of capturing video and audio data of live performances through cameras and microphones in the store;
[1722] A system including:
[1723] (Claim 2)
[1724] 2. The system according to claim 1, wherein the video data includes human movements, and the pre-processing means extracts poses and movement features of the human.
[1725] (Claim 3)
[1726] 2. The system of claim 1, wherein the audio data includes a musical performance, and the preprocessing means extracts musical features.
[1727] "Example 2: Combining Emotion Engines"
[1728] (Claim 1)
[1729] means for receiving video data or audio data;
[1730] means for pre-processing the received video or audio data;
[1731] means for analyzing the preprocessed data and calculating an art score;
[1732] a means of recognizing the emotions of competitors and spectators;
[1733] means for fine-tuning the artistic score based on the emotion recognition results;
[1734] a means for displaying the calculated artistic score;
[1735] A system including:
[1736] (Claim 2)
[1737] 2. The system according to claim 1, wherein the video data includes human movements, and the pre-processing means extracts poses and movement features of the human.
[1738] (Claim 3)
[1739] 2. The system of claim 1, wherein the audio data includes a musical performance, and the preprocessing means extracts musical features.
[1740] "Application example 2 when combining emotion engines"
[1741] (Claim 1)
[1742] means for receiving video data or audio data;
[1743] means for pre-processing the received video or audio data;
[1744] means for analyzing the preprocessed data and calculating an art score;
[1745] a means for displaying the calculated artistic score;
[1746] a means for recognizing a user's emotions and fine-tuning the rating based on the user's emotional response;
[1747] A system including:
[1748] (Claim 2)
[1749] 2. The system according to claim 1, wherein the video data includes human movements, and the pre-processing means extracts poses and movement features of the human.
[1750] (Claim 3)
[1751] 2. The system of claim 1, wherein the audio data includes a musical performance, and the preprocessing means extracts musical features.
[1752] (Claim 4)
[1753] 2. The system according to claim 1, wherein the means for recognizing the user's emotions analyzes the user's facial expressions and voice using an emotion engine.
[1754] (Claim 5)
[1755] 5. The system according to claim 4, further comprising means for collecting customer behavior and emotions using cameras and microphones placed in the store and providing the information to store employees. [Explanation of symbols]
[1756] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving video data or audio data; means for pre-processing the received video or audio data; means for analyzing the preprocessed data and calculating the artistic score; means for displaying the calculated artistic score; A system including:
2. 2. The system according to claim 1, wherein the video data includes human motion, and the pre-processing means extracts pose and motion characteristics of the human.
3. 2. The system of claim 1, wherein the audio data includes a musical performance, and the preprocessing means extracts musical features.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A