system

The system analyzes video and audio data to evaluate a musician's performance, calculating a comprehensive score for effective matching with suitable music group members, addressing the inefficiencies in conventional recruitment methods.

JP2026070983APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Conventional systems fail to accurately evaluate a musical instrument player's performance technique and musical taste, making it difficult to match them with appropriate music group members, leading to inefficient recruitment.

Method used

A system that analyzes video data to grasp the performer's movements and acoustics, calculates an overall musicality score, and recommends suitable music group members based on this evaluation.

Benefits of technology

Enables accurate and efficient matching of musicians based on their skills and musical preferences, improving the recruitment process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070983000001_ABST
    Figure 2026070983000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for receiving and saving video data recorded by a user during a performance, An analysis means for analyzing movements related to musical instrument performance from the video data, An acoustic analysis means for analyzing sound from the moving image data and evaluating its musicality, A method for calculating a comprehensive musicality score by integrating the results of video analysis and acoustic analysis, A means of comparing the user's desired musical style with the overall musicality score and recommending appropriate music group members, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , ,

[0005] , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the means for a conventional musical instrument player to search for members of a music group, there is a problem that the performance technique and musical taste online cannot be accurately evaluated, and it is difficult to match with appropriate members. In particular, there is a problem that efficient member recruitment cannot be achieved due to discrepancies in taste and inconsistencies in technical levels.

Means for Solving the Problems

[0005] This invention solves these problems by providing a system that analyzes video data recorded from a performance and comprehensively evaluates the movements and acoustics of the instrument performance. This system grasps the fingering and movement characteristics of the performer through video analysis and evaluates musical characteristics through audio analysis. It then integrates this data to calculate a comprehensive musicality score and recommends the most suitable music group members after comparing it with the musicality desired by the user. This enables the user to find members more accurately and effectively.

[0006] A "user" refers to an entity that uses the system to record their own performances and to recruit or match members for music groups.

[0007] "Video data" refers to data containing video and audio information of performances recorded by the user.

[0008] "Analysis means" refers to the technical means responsible for the process of analyzing the movements of a performer related to playing a musical instrument from video data.

[0009] "Acoustic analysis means" refers to technical means that are responsible for the process of analyzing audio information contained in video data and evaluating musical elements.

[0010] The "Overall Musicality Score" refers to a numerical index that evaluates a user's musical characteristics, calculated by integrating the results of video and audio analysis.

[0011] A "music group member" refers to a person with whom the user engages in musical activities that align with their desired musical style.

[0012] "Recommendation methods" refer to technical means that compare a user's overall musicality score with their desired musicality and suggest the most suitable music group members. [Brief explanation of the drawing]

[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system in which users upload videos of their performances to a platform, and an AI model is used to evaluate their performance skills and musical characteristics, and then recommend the most suitable music group members. The following describes an embodiment of this system.

[0035] User actions:

[0036] Users record their performances in video format using devices such as smartphones or PCs. They then upload the recorded videos to the system's dedicated platform. The videos include the performer's movements and audio information.

[0037] Server analysis processing:

[0038] The server first saves the uploaded video to storage. Next, it runs a video analysis AI model to analyze the saved video. The video analysis detects the user's movements and identifies skeletal movements and fingerings. The acoustic analysis module evaluates musical elements based on the audio extracted from the video. This includes variability in tempo, pitch, and volume.

[0039] Score calculation and matching:

[0040] The server integrates the results of motion and acoustic analysis to calculate an overall musicality score. This score quantifies the user's performance skills and musical expressiveness. Next, it compares the calculated score with the user's registered desired musicality. The server then filters the data to recommend suitable music group members to the user and generates a list.

[0041] Specific example:

[0042] For example, suppose a user uploads a video of themselves playing in a "jazz" style. The server analyzes the user's fingering and unique sonic nuances from this video and evaluates their overall musicality score as "Jazz Elegance." The system then selects members who are predicted to be a good match for a user with a "Jazz Elegance" score and presents them to the user as a list.

[0043] In this way, the system of the present invention can effectively match members based on the user's playing skills and musical preferences.

[0044] The following describes the processing flow.

[0045] Step 1:

[0046] Users record their own performances using devices such as smartphones or PCs, and save them as video files.

[0047] Step 2:

[0048] The device selects the video recorded by the user and opens the platform's dedicated upload screen. The user presses the upload button to send the video to the platform.

[0049] Step 3:

[0050] The server receives the uploaded video data and saves it to storage. This prepares it for subsequent analysis.

[0051] Step 4:

[0052] The server starts an AI video analysis model to analyze the stored video. This analysis detects the user's skeletal structure and finger movements from the video.

[0053] Step 5:

[0054] The server extracts audio data from the video and uses an acoustic analysis module to analyze tempo, pitch, and volume changes. This quantifies the musical elements.

[0055] Step 6:

[0056] The server integrates the results of video and audio analysis to calculate an overall musicality score. This score reflects the user's performance skills and musical expressiveness.

[0057] Step 7:

[0058] The server begins matching users with music group members based on their preferred musical style.

[0059] Step 8:

[0060] The server compares the calculated overall musicality score with the user's desired musicality and filters out the most suitable candidates for music group members.

[0061] Step 9:

[0062] The server generates a list of potential music group members based on the filtering results and sends it to the user's terminal.

[0063] Step 10:

[0064] The device displays a list of potential members to the user. The user then decides whether to contact members they are interested in.

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] In musical activities, finding suitable musical collaborators who match one's performance skills and musical characteristics is difficult. Therefore, there is a need for a method to efficiently evaluate a performer's abilities and musical characteristics and recommend the most suitable musical collaborators.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes means for receiving and storing multimedia data recorded by a user's performance, processing means for analyzing the performance actions from the multimedia data, and sound processing means for analyzing acoustic information from the multimedia data and evaluating musical characteristics. This makes it possible to accurately evaluate the performer's skill and musical characteristics and quickly recommend the most suitable musical collaborator to the user.

[0070] A "user" refers to an individual or group that uses the system to record their performances and receive evaluations and recommendations for their musicality.

[0071] "Multimedia data" refers to complex information data that includes audio and video, and specifically data that includes the actions of performers and the sounds of their performances.

[0072] "Performance actions" refer to the physical actions involved in playing a musical instrument, including body movements and finger movements.

[0073] "Processing means" refers to methods and techniques for analyzing and interpreting data in order to achieve a specific purpose.

[0074] "Acoustic information" refers to data that includes sound characteristics and elements related to music, such as tempo, pitch, and volume.

[0075] "Musical characteristics" refer to elements that indicate distinctive style and expressiveness in music, and serve as criteria for evaluating musicality.

[0076] "Acoustic processing means" refers to methods and equipment for effectively analyzing and evaluating acoustic information.

[0077] A "musical collaborator" refers to a partner or group member best suited to the user's performance skills and musical characteristics.

[0078] This invention is a system that allows users to record their own performances, evaluate their musical characteristics, and recommend appropriate musical collaborators. Specific embodiments of the invention are described below.

[0079] Users record their performances in video format using devices such as smartphones or PCs. Typically, high-resolution cameras and high-quality microphones are used for this purpose. The recorded video contains information about the performer's movements and audio, and this data is treated as multimedia data.

[0080] Users upload captured multimedia data to a dedicated platform on a server via an internet connection. The server acts as temporary storage for this multimedia data. Encryption technology is used on this platform to ensure data security.

[0081] The server uses a video analysis AI model to analyze the user's playing movements. This analysis identifies skeletal movements and finger placement, and recognizes movement patterns. The server also uses an acoustic analysis module as an acoustic processing tool to analyze acoustic information such as tempo, pitch, and volume fluctuations.

[0082] Based on the analysis, the server evaluates the user's musical characteristics and quantifies them as a musicality index. This musicality index is then compared to the user's desired musical characteristics, and a list of recommended musical collaborators is generated.

[0083] For example, if a user uploads a video of themselves playing a jazz-style instrument, the server analyzes the fingering and nuances of the sound and evaluates the musicality as "jazz elegance." Based on this, the server lists and recommends to the user musical collaborators who are a good match for the user and who share the "jazz elegance" characteristic.

[0084] An example of a prompt message is, "Evaluate the musicality of the jazz-style performance video and recommend suitable music group members." In this way, the system of the present invention can recommend effective musical collaborators based on the user's performance skills and musical characteristics.

[0085] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0086] Step 1:

[0087] Users record their performances in video format using their smartphones or PCs and upload the video data to the system's dedicated platform. The input is the performance video data, and the output is the video data stored on the platform. Specifically, the user presses the start recording button to begin playing, and then clicks the upload button to send the data after finishing.

[0088] Step 2:

[0089] The server stores the uploaded video data in temporary storage. The input is the video data received from the user, and the output is the file path in the storage. At this stage, the server checks the format and integrity of the video data and stores it in a secure environment. Specifically, it verifies the file format and writes the data to the storage.

[0090] Step 3:

[0091] The server uses a video analysis AI model to analyze performance movements from video data. The input is video data stored in storage, and the output is the analyzed movement data. Specifically, the AI ​​model recognizes the performer's skeletal structure and finger movements frame by frame and extracts movement patterns as digital data.

[0092] Step 4:

[0093] The server uses acoustic processing capabilities to analyze acoustic information from video data. The input is audio data stored in storage, and the output is acoustic characteristic data resulting from the analysis. Specifically, the acoustic processing module acquires information on tempo, pitch, and volume fluctuations and processes the data to evaluate musical characteristics.

[0094] Step 5:

[0095] The server integrates the analysis results of performance actions and acoustic characteristics to calculate an overall musicality index. The input consists of performance data and acoustic characteristic data, and the output is a musicality index score. Specifically, it compares each analysis result, weights them based on specific criteria, and generates an overall score.

[0096] Step 6:

[0097] The server generates a list of recommended musicians based on the calculated musicality index. The input is the musicality index score, and the output is a list of recommended musicians. Specifically, the server's filtering algorithm creates a list by matching it with the user's desired characteristics, and this list is displayed on the user's terminal.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] Individual musicians often find it difficult to accurately assess their own performance abilities and musical preferences. Furthermore, the lack of objective criteria for finding suitable musical group members results in a time-consuming and laborious process. This invention aims to provide a system that automatically analyzes a user's performance abilities and musical characteristics and quickly recommends suitable collaborators based on the results.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes a device for receiving and storing video data recorded by a user, a device for analyzing movements related to the performance from the video data, and a device for analyzing sound from the video data and evaluating musicality. This makes it possible to efficiently analyze the user's performance ability and musical characteristics and recommend appropriate collaborators.

[0103] A "user" is the entity that records video and image data and uploads it to the system.

[0104] "Performing" is the act of expressing music using a musical instrument.

[0105] "Video data" refers to data in video format that records a performance.

[0106] A "storage device" is a mechanism for storing received video and image data.

[0107] An "analysis device" is a technology for detecting and interpreting movements related to a performance from video data.

[0108] A "device for evaluating musicality" is a means of analyzing sound and quantifying the technical and expressive characteristics of a performance.

[0109] "Musicality evaluation" is an indicator of overall performance ability obtained through video data and acoustic analysis.

[0110] A "collaborator" is a member who can work together on musical activities.

[0111] A "recommended device" is a system that suggests appropriate collaborators based on the user's musicality evaluation.

[0112] The "presentation device" is an interface for informing the user of recommended collaborators.

[0113] The system for implementing this invention consists of a user recording their performance using a portable information terminal such as a smartphone and uploading the video data to a server. The server first stores the received video data. Then, it uses a cloud-based analysis service to analyze the performer's movements and music data in the video. This analysis includes a video analysis device and specific software modules for evaluating musicality.

[0114] The server uses video analysis software such as Google Cloud Video Intelligence and AWS Rekognition to analyze the performer's technique and physical structure from video data. Furthermore, it uses Google Cloud Speech-to-Text and Amazon Transcribe to extract musical characteristics from audio data and evaluate musicality. These results are integrated to generate a comprehensive musicality evaluation.

[0115] Subsequently, the server uses a generative AI model to extract the most suitable collaborators based on musicality evaluations and generates a user-specific recommendation list. This process includes a scoring algorithm using a programming language such as Python. On the user's device, an interface is built to receive and display the recommendation list from the server.

[0116] For example, if a user uploads a video of themselves playing classical piano, the server will analyze it, evaluate the user's playing technique according to classical music standards, and recommend other performers with similar musicality. This analysis can be performed using a prompt message such as, "This video contains a performance of classical music. Please analyze the fingering accuracy and musical expressiveness."

[0117] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0118] Step 1:

[0119] Users record their performances as videos on devices such as smartphones and upload them to the server using the system's application. The input is the user's performance video file, and the output is the video data stored on the server.

[0120] Step 2:

[0121] The server passes the received video data to video analysis software (e.g., Google Cloud Video Intelligence) to analyze the performer's technique and physical structure. In this step, the input is video data, and the output is analysis information (e.g., finger movements and rhythm patterns).

[0122] Step 3:

[0123] The server extracts audio data and performs acoustic analysis. The audio data is then fed into Google Cloud Speech-to-Text or Amazon Transcribe to extract musical features (e.g., tempo, pitch). The input is audio data, and the output is musical feature data.

[0124] Step 4:

[0125] The server integrates data obtained from video analysis software and audio analysis software to generate a comprehensive musicality evaluation. It uses video analysis information and musical characteristic data as input, and outputs a musicality evaluation score.

[0126] Step 5:

[0127] Using a generative AI model, the server lists suitable collaborators based on musicality evaluation scores. In this step, the input is the musicality evaluation score, and the output is a recommendation list. The prompt message can be "This video contains a performance of classical music. Please analyze the fingering accuracy and musical expressiveness."

[0128] Step 6:

[0129] The server sends the generated recommendation list to the user's terminal, which then launches an interface to display the list. The input is the recommendation list, and the output is a list of collaborators displayed on the terminal's screen.

[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0131] This invention combines an emotion engine with a system that comprehensively analyzes data obtained from users' performance videos and recommends appropriate members for a music group, thereby achieving more personalized matching that takes into account the user's emotional state. The embodiments of this system are described in detail below.

[0132] User actions:

[0133] Users record their performances using devices such as smartphones or PCs. The recordings include both video and audio of the performance, and users upload these video files to the platform.

[0134] Server processing:

[0135] The server first saves the received video data to storage. For the video, an AI model for video analysis is used to perform skeletal detection and finger movement analysis, and for the audio, an acoustic analysis module is used to evaluate the characteristics of the music. In this process, an emotion engine is also used to analyze the user's facial expressions and movements from the video and obtain emotion data.

[0136] Score calculation and adjustment:

[0137] The server integrates video analysis, acoustic analysis, and emotional data to calculate an overall musicality score. Based on the emotional data obtained by the emotion engine, the score can be adjusted to take into account the user's emotional state.

[0138] Matching process:

[0139] The server uses the calculated overall musicality score to match the user with their desired musical style. Based on emotional data, it further refines the filtering process to generate a list of suitable music group members for the user.

[0140] Specific example:

[0141] For example, if a user uploads a video of themselves playing the blues, the emotion engine might recognize that the user experienced high levels of satisfaction during their performance. The server would then consider this positive emotion data and list other members who share similar emotional expressiveness with other members specializing in the blues. This emotional data-driven matching process can foster richer musical interaction between the user and existing members.

[0142] This system enables personalized music group matching that takes into account the user's emotional state, in addition to traditional technical and musical evaluations.

[0143] The following describes the processing flow.

[0144] Step 1:

[0145] Users record their performances as videos on their smartphones or PCs. These videos should capture the performer's facial expressions and movements.

[0146] Step 2:

[0147] The device sends the recorded video file to the server by opening the platform's dedicated upload screen, selecting the file, and pressing the upload button.

[0148] Step 3:

[0149] The server receives the uploaded video data and saves it to storage. This storage enables subsequent analysis processes.

[0150] Step 4:

[0151] The server inputs the saved videos into an AI video analysis model to analyze the performer's skeletal structure and finger movements. This allows for the quantification of the characteristics of the performance technique.

[0152] Step 5:

[0153] The server passes the audio data extracted from the video to an audio analysis module, which analyzes the tempo, pitch, and volume changes of the music. The musical characteristics are then quantified.

[0154] Step 6:

[0155] The server activates an emotion engine and analyzes the user's facial expressions and gestures in the video. This allows the emotional state during the performance to be quantified.

[0156] Step 7:

[0157] The server integrates the results of video analysis, audio analysis, and emotion engine analysis to calculate an overall musicality score. The score is then adjusted based on the recognized emotion data.

[0158] Step 8:

[0159] The server compares the user's registered musical preferences with an overall musicality score to search for suitable music group members. It further narrows down the candidates by utilizing emotional data.

[0160] Step 9:

[0161] The server generates a list of optimized music group member candidates and sends it to the user's terminal.

[0162] Step 10:

[0163] The device displays a list of potential members to the user. The user then decides whether to begin interacting with members from that list who interest them.

[0164] (Example 2)

[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0166] In musical activities, for performers to collaborate with other members, not only technical ability but also musicality and emotional compatibility are important. However, conventional systems do not take such emotional compatibility into consideration when making recommendations, which can lead to inappropriate matches. Therefore, in order to improve users' satisfaction with their musical activities, there is a need for recommendations of music group members that take emotional aspects into account.

[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0168] In this invention, the server includes means for receiving and storing video footage recorded by the user, means for analyzing actions related to musical instrument performance from the video footage, and acoustic analysis means for analyzing sound from the video footage and evaluating musicality. This makes it possible to adjust the overall musicality evaluation based on emotional data and recommend members of a musical group that take into account the user's emotional state.

[0169] A "user" is the entity that uses the system to record a performance and uploads that data.

[0170] "Video and audio materials" refers to video files recorded by users during performances, and includes both video and audio.

[0171] "Skeleton" refers to a data structure that shows the location of the main bones and joints of a performer's body within video footage.

[0172] "Emotional data" refers to information indicating the emotional state extracted from the user's facial expressions and movements, and is used to evaluate emotional expression during performance.

[0173] "Comprehensive Musical Quality Evaluation" is an evaluation index that shows musical quality by integrating video analysis results, sound analysis results, and emotional data.

[0174] "Music group members" refers to other members who are encouraged to collaborate with the user in musical activities.

[0175] The "list" is a compilation of candidate members for a music group, selected based on the user's playing style and emotional state.

[0176] A "server" is a computer device that serves as the core of the entire system, responsible for data storage, analysis, and the generation of recommendation results.

[0177] This system uses video recordings of user performances to appropriately match users with other members of musical groups. Users record their performances using devices such as smartphones or PCs and upload the video data to the platform. The recorded videos include both video and audio of the performance. The system aims to recommend the most suitable members for a musical group based on the user's performance style and emotions.

[0178] The server is equipped with an AI model that analyzes video footage and evaluates the user's movements and musical characteristics during performance. First, the server saves the uploaded video data to its storage. Next, it uses the video analysis AI model to perform skeletal detection and finger movement analysis of the user. This analysis allows for a precise understanding of the user's movements related to playing the instrument.

[0179] The acoustic analysis module analyzes audio data and evaluates the tempo, rhythm, and musicality of the music. Based on these analysis results, the system identifies the user's musical genre and playing style.

[0180] Furthermore, an emotion engine is used to acquire emotional data from the user's facial expressions and actions within the video. This makes it possible to adjust the overall musicality evaluation while considering the user's emotional state. The emotional data reflects the user's emotional expression during performance and is used to filter out appropriate musical components.

[0181] As a concrete example, consider a scenario where a user uploads a video of themselves playing "jazz," and the emotion engine recognizes high levels of excitement and pride in the user's performance. Based on this information, the server creates a list of members with similar emotional expressiveness specializing in jazz performance. This process facilitates rich musical interaction based on emotional data.

[0182] An example of a prompt message is: "We will analyze the subject's face and movements from this user's performance video, identify their emotional state during the performance, and use that data to adjust the overall musicality evaluation. Then, based on the adjusted evaluation, we will recommend suitable music group members."

[0183] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0184] Step 1:

[0185] Users record their performances using devices such as smartphones or PCs and upload the video data to the system's platform. The input here is a video file containing both video and audio. This file forms the basis for subsequent analysis steps.

[0186] Step 2:

[0187] The server receives video files uploaded by users and saves them to storage. During this saving process, the server checks the format of the video files and converts them to a format suitable for analysis if necessary. This conversion ensures that the server has access to the data needed for analysis.

[0188] Step 3:

[0189] The server activates a video analysis AI model to analyze the video portion of the footage. Here, a skeletal detection algorithm is used to identify the movements of the performer's body and fingers, and fingering data is extracted. The input is video data, and the output is quantitative data indicating the performer's movements.

[0190] Step 4:

[0191] The server uses an acoustic analysis module to analyze the audio portion of a video and evaluate its musical characteristics, such as tempo, rhythm, and genre. The input for this process is audio data obtained from the video, and the output is data that describes the musical characteristics in detail.

[0192] Step 5:

[0193] The server uses an emotion engine to acquire emotional data based on the user's facial expressions and actions in the video. The input is video data, and the output is qualitative data indicating the user's emotional state. This information is used to understand emotional expression during performance.

[0194] Step 6:

[0195] The server integrates the previously obtained video analysis results, acoustic analysis results, and emotion data to calculate an overall musicality rating. In this integration process, individual data points are converted into evaluation scales that indicate the user's musical skills and emotional expressiveness. The output is an overall musicality score.

[0196] Step 7:

[0197] The server performs a process of recommending the most suitable members of a musical group to the user based on an overall musicality evaluation. It also considers emotional compatibility based on emotional data, filters the recommendations, generates a final recommendation list, and sends it to the user's terminal. The input is the overall musicality evaluation, and the output is a list of recommended musical members.

[0198] (Application Example 2)

[0199] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0200] Traditional music content recommendation systems primarily rely on recommendations based on musical skill and preferences, making it difficult to provide personalized recommendations that reflect the listener's emotional state. Furthermore, collaborator recommendations, designed to help users maximize their musical potential, also lack emotional considerations for personalization. Therefore, there are limitations to providing the optimal music experience for listeners, and a challenge remains in enhancing users' emotional satisfaction with their musical activities.

[0201] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0202] In this invention, the server includes means for receiving and storing visual and auditory data recorded by the user; analysis means for analyzing movements related to musical instrument performance from the visual and auditory data; acoustic analysis means for analyzing sound from the visual and auditory data and evaluating musicality; means for analyzing emotional state from the visual and auditory data and acquiring emotional information; means for integrating the results of the video analysis, acoustic analysis, and emotional information to calculate an overall musicality score; and means for comparing the user's desired musicality with the overall musicality score and recommending an appropriate musical collaborator. This makes it possible to recommend music content according to the viewer's emotional state and to provide personalized recommendations for the most suitable musical collaborator for the user.

[0203] "Visual and auditory data" refers to video and audio information recorded by the user, and serves as basic data for analyzing movements and musicality related to musical instrument performance.

[0204] "Analysis means" refers to a device or software for extracting movements and characteristics related to musical instrument performance from visual and auditory data and analyzing that information.

[0205] "Acoustic analysis means" refers to a device or software for evaluating sound characteristics from visual and auditory data and determining musicality.

[0206] "Emotional information" refers to data that indicates the emotional state expressed through the user's facial expressions and actions, obtained from visual and auditory data.

[0207] The "Comprehensive Musicality Score" is a numerical value or index that represents a comprehensive evaluation that integrates the results of visual and auditory data analysis and includes musical characteristics and emotional information.

[0208] A "music collaborator" is a partner who collaborates with a user on musical activities, recommended based on emotional information and musicality.

[0209] This system consists of client-side smart devices and server-side data processing facilities. The smart devices are responsible for recording the user's performance as visual and auditory data. Specifically, smartphones and tablets are used, and their cameras and microphones are used to acquire video and audio. The acquired data is transmitted to the server via the internet.

[0210] The server performs various processes using the received visual and auditory data. First, it analyzes the visual data using image processing libraries such as OpenCV to detect playing movements and skeletal movements. This information is used to extract the technical characteristics of the instrument performance. Meanwhile, it utilizes acoustic analysis libraries such as Librosa to analyze the musicality of the acoustic data. This allows for the evaluation of various musical characteristics such as the tempo, rhythm, and melody line of the piece.

[0211] Furthermore, using emotion analysis engines such as Hume AI, the system analyzes the user's facial expressions from visual data to obtain emotional information. This emotional information reflects the user's emotional state during performance and deeply influences the overall musicality score generated by the server. This makes it possible to score not only based on technical skills but also on emotional expressiveness.

[0212] The server matches the user's desired musical style with their overall musicality score and lists suitable musical collaborators. This list is displayed on the user's device, allowing the user to select the collaborator best suited to their musical activities.

[0213] For example, suppose a user plays jazz piano and expresses a sense of exhilaration in their performance. In this case, the system would recommend a musical collaborator who also has an interest in jazz and can play with a similar sense of exhilaration. Another example of a prompt in a generative AI model could be: "Based on the emotional state detected from the user's performance video, recommend jazz content that will give viewers a sense of exhilaration."

[0214] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0215] Step 1: The user records their performance using a smart device and saves it as visual and auditory data. After recording, the device sends this data to a server via the internet. In this step, the video and audio of the performance are inputs, and the data reaches the server as output.

[0216] Step 2: The server analyzes the received visual data using the OpenCV library. This involves detecting the performer's skeletal structure and movements from the video data and obtaining data to evaluate their performance technique. The input is visual data, and the output is the analyzed motion information.

[0217] Step 3: The server analyzes the auditory data using the Librosa library. It extracts musical features, such as tempo and pitch, from the audio data and generates data to evaluate musicality based on these features. The input is auditory data, and the output is the analyzed musical features.

[0218] Step 4: The server uses an emotion analysis engine, such as Hume AI, to analyze visual data and obtain information about the user's emotions. Specifically, it identifies the emotional state during performance from facial expressions and body movements. The input is visual data, and the output is data indicating the emotional state.

[0219] Step 5: The server integrates the previously acquired behavioral information, musical characteristics, and emotional information to calculate an overall musicality score. This score reflects the user's musical skills and emotional expression abilities. The input is the data obtained in Steps 2 through 4, and the output is the overall musicality score.

[0220] Step 6: The server compares the user's desired musicality with the overall musicality score and generates a list to recommend the most suitable musical collaborators. The inputs are the overall musicality score and the user's desired musicality, and the output is a list of collaborators. This list is provided to the user's terminal, enabling matching with appropriate collaborators.

[0221] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0222] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0223] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0224] [Second Embodiment]

[0225] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0226] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0227] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0228] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0229] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0230] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0231] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0232] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0233] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0234] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0235] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0236] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0237] This invention is a system in which users upload videos of their performances to a platform, and an AI model is used to evaluate their performance skills and musical characteristics, and then recommend the most suitable music group members. The following describes an embodiment of this system.

[0238] User actions:

[0239] Users record their performances in video format using devices such as smartphones or PCs. They then upload the recorded videos to the system's dedicated platform. The videos include the performer's movements and audio information.

[0240] Server analysis processing:

[0241] The server first saves the uploaded video to storage. Next, it runs a video analysis AI model to analyze the saved video. The video analysis detects the user's movements and identifies skeletal movements and fingerings. The acoustic analysis module evaluates musical elements based on the audio extracted from the video. This includes variability in tempo, pitch, and volume.

[0242] Score calculation and matching:

[0243] The server integrates the results of motion and acoustic analysis to calculate an overall musicality score. This score quantifies the user's performance skills and musical expressiveness. Next, it compares the calculated score with the user's registered desired musicality. The server then filters the data to recommend suitable music group members to the user and generates a list.

[0244] Specific example:

[0245] For example, suppose a user uploads a video of themselves playing in a "jazz" style. The server analyzes the user's fingering and unique sonic nuances from this video and evaluates their overall musicality score as "Jazz Elegance." The system then selects members who are predicted to be a good match for a user with a "Jazz Elegance" score and presents them to the user as a list.

[0246] In this way, the system of the present invention can effectively match members based on the user's playing skills and musical preferences.

[0247] The following describes the processing flow.

[0248] Step 1:

[0249] Users record their own performances using devices such as smartphones or PCs, and save them as video files.

[0250] Step 2:

[0251] The device selects the video recorded by the user and opens the platform's dedicated upload screen. The user presses the upload button to send the video to the platform.

[0252] Step 3:

[0253] The server receives the uploaded video data and saves it to storage. This prepares it for subsequent analysis.

[0254] Step 4:

[0255] The server starts an AI video analysis model to analyze the stored video. This analysis detects the user's skeletal structure and finger movements from the video.

[0256] Step 5:

[0257] The server extracts audio data from the video and uses an acoustic analysis module to analyze tempo, pitch, and volume changes. This quantifies the musical elements.

[0258] Step 6:

[0259] The server integrates the results of video and audio analysis to calculate an overall musicality score. This score reflects the user's performance skills and musical expressiveness.

[0260] Step 7:

[0261] The server begins matching users with music group members based on their preferred musical style.

[0262] Step 8:

[0263] The server compares the calculated overall musicality score with the user's desired musicality and filters out the most suitable candidates for music group members.

[0264] Step 9:

[0265] The server generates a list of potential music group members based on the filtering results and sends it to the user's terminal.

[0266] Step 10:

[0267] The device displays a list of potential members to the user. The user then decides whether to contact members they are interested in.

[0268] (Example 1)

[0269] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0270] In musical activities, finding suitable musical collaborators who match one's performance skills and musical characteristics is difficult. Therefore, there is a need for a method to efficiently evaluate a performer's abilities and musical characteristics and recommend the most suitable musical collaborators.

[0271] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0272] In this invention, the server includes means for receiving and storing multimedia data recorded by a user's performance, processing means for analyzing the performance actions from the multimedia data, and sound processing means for analyzing acoustic information from the multimedia data and evaluating musical characteristics. This makes it possible to accurately evaluate the performer's skill and musical characteristics and quickly recommend the most suitable musical collaborator to the user.

[0273] A "user" refers to an individual or group that uses the system to record their performances and receive evaluations and recommendations for their musicality.

[0274] "Multimedia data" refers to complex information data that includes audio and video, and specifically data that includes the actions of performers and the sounds of their performances.

[0275] "Performance actions" refer to the physical actions involved in playing a musical instrument, including body movements and finger movements.

[0276] "Processing means" refers to methods and techniques for analyzing and interpreting data in order to achieve a specific purpose.

[0277] "Acoustic information" refers to data that includes sound characteristics and elements related to music, such as tempo, pitch, and volume.

[0278] "Musical characteristics" refer to elements that indicate distinctive style and expressiveness in music, and serve as criteria for evaluating musicality.

[0279] "Acoustic processing means" refers to methods and equipment for effectively analyzing and evaluating acoustic information.

[0280] A "musical collaborator" refers to a partner or group member best suited to the user's performance skills and musical characteristics.

[0281] This invention is a system that allows users to record their own performances, evaluate their musical characteristics, and recommend appropriate musical collaborators. Specific embodiments of the invention are described below.

[0282] Users record their performances in video format using devices such as smartphones or PCs. Typically, high-resolution cameras and high-quality microphones are used for this purpose. The recorded video contains information about the performer's movements and audio, and this data is treated as multimedia data.

[0283] The user uploads the captured multimedia data to a dedicated platform on the server via an Internet connection. The server serves as storage to temporarily store this multimedia data. Encryption technology is used on this platform to ensure data security.

[0284] The server uses a video analysis AI model to analyze the user's performance actions. In this analysis, the movement of the skeleton and finger gestures are identified, and the pattern of the actions is recognized. Also, the server utilizes an acoustic analysis module as an acoustic processing means to analyze acoustic information such as tempo, pitch, and volume variation.

[0285] As a result of the analysis, the server evaluates the user's music characteristics and quantifies them as music property indicators. These music property indicators are compared with the music characteristics desired by the user, and a recommended list of optimal music collaborators is generated.

[0286] For example, when the user uploads a jazz-style performance video, the server analyzes the fingering and nuances of the sound of the instrument and evaluates the music property indicator as "jazz elegance". Based on this, the server lists up music collaborators who are compatible with other users having the characteristics of "jazz elegance" and recommends them to the user.

[0287] Examples of the prompt sentence include "Evaluate the musicality from a jazz-style performance video and recommend suitable music group members." In this way, the system of the present invention can effectively recommend music collaborators based on the user's performance technique and music characteristics.

[0288] The flow of the specific process in Example 1 will be described using FIG. 11.

[0289] Step 1:

[0290] Users record their performances in video format using their smartphones or PCs and upload the video data to the system's dedicated platform. The input is the performance video data, and the output is the video data stored on the platform. Specifically, the user presses the start recording button to begin playing, and then clicks the upload button to send the data after finishing.

[0291] Step 2:

[0292] The server stores the uploaded video data in temporary storage. The input is the video data received from the user, and the output is the file path in the storage. At this stage, the server checks the format and integrity of the video data and stores it in a secure environment. Specifically, it verifies the file format and writes the data to the storage.

[0293] Step 3:

[0294] The server uses a video analysis AI model to analyze performance movements from video data. The input is video data stored in storage, and the output is the analyzed movement data. Specifically, the AI ​​model recognizes the performer's skeletal structure and finger movements frame by frame and extracts movement patterns as digital data.

[0295] Step 4:

[0296] The server uses acoustic processing capabilities to analyze acoustic information from video data. The input is audio data stored in storage, and the output is acoustic characteristic data resulting from the analysis. Specifically, the acoustic processing module acquires information on tempo, pitch, and volume fluctuations and processes the data to evaluate musical characteristics.

[0297] Step 5:

[0298] The server integrates the analysis results of performance actions and acoustic characteristics to calculate an overall musicality index. The input consists of performance data and acoustic characteristic data, and the output is a musicality index score. Specifically, it compares each analysis result, weights them based on specific criteria, and generates an overall score.

[0299] Step 6:

[0300] The server generates a list of recommended musicians based on the calculated musicality index. The input is the musicality index score, and the output is a list of recommended musicians. Specifically, the server's filtering algorithm creates a list by matching it with the user's desired characteristics, and this list is displayed on the user's terminal.

[0301] (Application Example 1)

[0302] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0303] Individual musicians often find it difficult to accurately assess their own performance abilities and musical preferences. Furthermore, the lack of objective criteria for finding suitable musical group members results in a time-consuming and laborious process. This invention aims to provide a system that automatically analyzes a user's performance abilities and musical characteristics and quickly recommends suitable collaborators based on the results.

[0304] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0305] In this invention, the server includes a device for receiving and storing video data recorded by a user, a device for analyzing movements related to the performance from the video data, and a device for analyzing sound from the video data and evaluating musicality. This makes it possible to efficiently analyze the user's performance ability and musical characteristics and recommend appropriate collaborators.

[0306] A "user" is the entity that records moving image data and uploads it to the system.

[0307] A "performance" is an act of expressing music using musical instruments.

[0308] "Moving image data" is data in a video format that records the state of a performance.

[0309] A "storage device" is a mechanism for holding the received moving image data.

[0310] An "analysis device" is a technology for detecting and interpreting movements related to a performance from moving image data.

[0311] A "device for evaluating musicality" is a means for analyzing sounds and quantifying the technical and expressive characteristics of a performance.

[0312] "Musicality evaluation" is an index of the overall performance ability obtained from moving image data and acoustic analysis.

[0313] A "collaborator" is a member who can collaborate in musical activities.

[0314] A "recommendation device" is a system for presenting appropriate collaborators based on the musicality evaluation of a user.

[0315] A "presentation device" is an interface for conveying the recommended collaborators to the user.

[0316] The system for implementing this invention consists of a user recording their performance using a portable information terminal such as a smartphone and uploading the video data to a server. The server first stores the received video data. Then, it uses a cloud-based analysis service to analyze the performer's movements and music data in the video. This analysis includes a video analysis device and specific software modules for evaluating musicality.

[0317] The server uses video analysis software such as Google Cloud Video Intelligence and AWS Rekognition to analyze the performer's technique and physical structure from video data. Furthermore, it extracts musical characteristics from audio data using Google Cloud Speech-to-Text and Amazon Transcribe to evaluate musicality. These results are integrated to generate a comprehensive musicality evaluation.

[0318] Subsequently, the server uses a generative AI model to extract the most suitable collaborators based on musicality evaluations and generates a user-specific recommendation list. This process includes a scoring algorithm using a programming language such as Python. On the user's device, an interface is built to receive and display the recommendation list from the server.

[0319] For example, if a user uploads a video of themselves playing classical piano, the server will analyze it, evaluate the user's playing technique according to classical music standards, and recommend other performers with similar musicality. This analysis can be performed using a prompt message such as, "This video contains a performance of classical music. Please analyze the fingering accuracy and musical expressiveness."

[0320] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0321] Step 1:

[0322] Users record their performances as videos on devices such as smartphones and upload them to the server using the system's application. The input is the user's performance video file, and the output is the video data stored on the server.

[0323] Step 2:

[0324] The server passes the received video data to video analysis software (e.g., Google Cloud Video Intelligence) to analyze the performer's technique and physical structure. In this step, the input is video data, and the output is analysis information (e.g., finger movements and rhythm patterns).

[0325] Step 3:

[0326] The server extracts audio data and performs acoustic analysis. The audio data is then fed into Google Cloud Speech-to-Text or Amazon Transcribe to extract musical features (e.g., tempo, pitch). The input is audio data, and the output is musical feature data.

[0327] Step 4:

[0328] The server integrates data obtained from video analysis software and audio analysis software to generate a comprehensive musicality evaluation. It uses video analysis information and musical characteristic data as input, and outputs a musicality evaluation score.

[0329] Step 5:

[0330] Using a generative AI model, the server lists suitable collaborators based on musicality evaluation scores. In this step, the input is the musicality evaluation score, and the output is a recommendation list. The prompt message can be "This video contains a performance of classical music. Please analyze the fingering accuracy and musical expressiveness."

[0331] Step 6:

[0332] The server sends the generated recommendation list to the user's terminal, which then launches an interface to display the list. The input is the recommendation list, and the output is a list of collaborators displayed on the terminal's screen.

[0333] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0334] This invention combines an emotion engine with a system that comprehensively analyzes data obtained from users' performance videos and recommends appropriate members for a music group, thereby achieving more personalized matching that takes into account the user's emotional state. The embodiments of this system are described in detail below.

[0335] User actions:

[0336] Users record their performances using devices such as smartphones or PCs. The recordings include both video and audio of the performance, and users upload these video files to the platform.

[0337] Server processing:

[0338] The server first saves the received video data to storage. For the video, an AI model for video analysis is used to perform skeletal detection and finger movement analysis, and for the audio, an acoustic analysis module is used to evaluate the characteristics of the music. In this process, an emotion engine is also used to analyze the user's facial expressions and movements from the video and obtain emotion data.

[0339] Score calculation and adjustment:

[0340] The server integrates video analysis, acoustic analysis, and emotional data to calculate an overall musicality score. Based on the emotional data obtained by the emotion engine, the score can be adjusted to take into account the user's emotional state.

[0341] Matching process:

[0342] The server uses the calculated overall musicality score to match the user with their desired musical style. Based on emotional data, it further refines the filtering process to generate a list of suitable music group members for the user.

[0343] Specific example:

[0344] For example, if a user uploads a video of themselves playing the blues, the emotion engine might recognize that the user experienced high levels of satisfaction during their performance. The server would then consider this positive emotion data and list other members who share similar emotional expressiveness with other members specializing in the blues. This emotional data-driven matching process can foster richer musical interaction between the user and existing members.

[0345] This system enables personalized music group matching that takes into account the user's emotional state, in addition to traditional technical and musical evaluations.

[0346] The following describes the processing flow.

[0347] Step 1:

[0348] Users record their performances as videos on their smartphones or PCs. These videos should capture the performer's facial expressions and movements.

[0349] Step 2:

[0350] The device sends the recorded video file to the server by opening the platform's dedicated upload screen, selecting the file, and pressing the upload button.

[0351] Step 3:

[0352] The server receives the uploaded video data and saves it to storage. This storage enables subsequent analysis processes.

[0353] Step 4:

[0354] The server inputs the saved videos into an AI video analysis model to analyze the performer's skeletal structure and finger movements. This allows for the quantification of the characteristics of the performance technique.

[0355] Step 5:

[0356] The server passes the audio data extracted from the video to an audio analysis module, which analyzes the tempo, pitch, and volume changes of the music. The musical characteristics are then quantified.

[0357] Step 6:

[0358] The server activates an emotion engine and analyzes the user's facial expressions and gestures in the video. This allows the emotional state during the performance to be quantified.

[0359] Step 7:

[0360] The server integrates the results of video analysis, audio analysis, and emotion engine analysis to calculate an overall musicality score. The score is then adjusted based on the recognized emotion data.

[0361] Step 8:

[0362] The server compares the user's registered musical preferences with an overall musicality score to search for suitable music group members. It further narrows down the candidates by utilizing emotional data.

[0363] Step 9:

[0364] The server generates a list of optimized music group member candidates and sends it to the user's terminal.

[0365] Step 10:

[0366] The device displays a list of potential members to the user. The user then decides whether to begin interacting with members from that list who interest them.

[0367] (Example 2)

[0368] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0369] In musical activities, for performers to collaborate with other members, not only technical ability but also musicality and emotional compatibility are important. However, conventional systems do not take such emotional compatibility into consideration when making recommendations, which can lead to inappropriate matches. Therefore, in order to improve users' satisfaction with their musical activities, there is a need for recommendations of music group members that take emotional aspects into account.

[0370] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0371] In this invention, the server includes means for receiving and storing video footage recorded by the user, means for analyzing actions related to musical instrument performance from the video footage, and acoustic analysis means for analyzing sound from the video footage and evaluating musicality. This makes it possible to adjust the overall musicality evaluation based on emotional data and recommend members of a musical group that take into account the user's emotional state.

[0372] A "user" is the entity that uses the system to record a performance and uploads that data.

[0373] "Video and audio materials" refers to video files recorded by users during performances, and includes both video and audio.

[0374] "Skeleton" refers to a data structure that shows the location of the main bones and joints of a performer's body within video footage.

[0375] "Emotional data" refers to information indicating the emotional state extracted from the user's facial expressions and movements, and is used to evaluate emotional expression during performance.

[0376] "Comprehensive Musical Quality Evaluation" is an evaluation index that shows musical quality by integrating video analysis results, sound analysis results, and emotional data.

[0377] "Music group members" refers to other members who are encouraged to collaborate with the user in musical activities.

[0378] The "list" is a compilation of candidate members for a music group, selected based on the user's playing style and emotional state.

[0379] A "server" is a computer device that serves as the core of the entire system, responsible for data storage, analysis, and the generation of recommendation results.

[0380] This system uses video recordings of user performances to appropriately match users with other members of musical groups. Users record their performances using devices such as smartphones or PCs and upload the video data to the platform. The recorded videos include both video and audio of the performance. The system aims to recommend the most suitable members for a musical group based on the user's performance style and emotions.

[0381] The server is equipped with an AI model that analyzes video footage and evaluates the user's movements and musical characteristics during performance. First, the server saves the uploaded video data to its storage. Next, it uses the video analysis AI model to perform skeletal detection and finger movement analysis of the user. This analysis allows for a precise understanding of the user's movements related to playing the instrument.

[0382] The acoustic analysis module analyzes audio data and evaluates the tempo, rhythm, and musicality of the music. Based on these analysis results, the system identifies the user's musical genre and playing style.

[0383] Furthermore, an emotion engine is used to acquire emotional data from the user's facial expressions and actions within the video. This makes it possible to adjust the overall musicality evaluation while considering the user's emotional state. The emotional data reflects the user's emotional expression during performance and is used to filter out appropriate musical components.

[0384] As a concrete example, consider a scenario where a user uploads a video of themselves playing "jazz," and the emotion engine recognizes high levels of excitement and pride in the user's performance. Based on this information, the server creates a list of members with similar emotional expressiveness specializing in jazz performance. This process facilitates rich musical interaction based on emotional data.

[0385] An example of a prompt message is: "We will analyze the subject's face and movements from this user's performance video, identify their emotional state during the performance, and use that data to adjust the overall musicality evaluation. Then, based on the adjusted evaluation, we will recommend suitable music group members."

[0386] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0387] Step 1:

[0388] Users record their performances using devices such as smartphones or PCs and upload the video data to the system's platform. The input here is a video file containing both video and audio. This file forms the basis for subsequent analysis steps.

[0389] Step 2:

[0390] The server receives video files uploaded by users and saves them to storage. During this saving process, the server checks the format of the video files and converts them to a format suitable for analysis if necessary. This conversion ensures that the server has access to the data needed for analysis.

[0391] Step 3:

[0392] The server activates a video analysis AI model to analyze the video portion of the footage. Here, a skeletal detection algorithm is used to identify the movements of the performer's body and fingers, and fingering data is extracted. The input is video data, and the output is quantitative data indicating the performer's movements.

[0393] Step 4:

[0394] The server uses an acoustic analysis module to analyze the audio portion of a video and evaluate its musical characteristics, such as tempo, rhythm, and genre. The input for this process is audio data obtained from the video, and the output is data that describes the musical characteristics in detail.

[0395] Step 5:

[0396] The server uses an emotion engine to acquire emotional data based on the user's facial expressions and actions in the video. The input is video data, and the output is qualitative data indicating the user's emotional state. This information is used to understand emotional expression during performance.

[0397] Step 6:

[0398] The server integrates the previously obtained video analysis results, acoustic analysis results, and emotion data to calculate an overall musicality rating. In this integration process, individual data points are converted into evaluation scales that indicate the user's musical skills and emotional expressiveness. The output is an overall musicality score.

[0399] Step 7:

[0400] The server performs a process of recommending the most suitable members of a musical group to the user based on an overall musicality evaluation. It also considers emotional compatibility based on emotional data, filters the recommendations, generates a final recommendation list, and sends it to the user's terminal. The input is the overall musicality evaluation, and the output is a list of recommended musical members.

[0401] (Application Example 2)

[0402] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0403] Traditional music content recommendation systems primarily rely on recommendations based on musical skill and preferences, making it difficult to provide personalized recommendations that reflect the listener's emotional state. Furthermore, collaborator recommendations, designed to help users maximize their musical potential, also lack emotional considerations for personalization. Therefore, there are limitations to providing the optimal music experience for listeners, and a challenge remains in enhancing users' emotional satisfaction with their musical activities.

[0404] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0405] In this invention, the server includes means for receiving and storing visual and auditory data recorded by the user; analysis means for analyzing movements related to musical instrument performance from the visual and auditory data; acoustic analysis means for analyzing sound from the visual and auditory data and evaluating musicality; means for analyzing emotional state from the visual and auditory data and acquiring emotional information; means for integrating the results of the video analysis, acoustic analysis, and emotional information to calculate an overall musicality score; and means for comparing the user's desired musicality with the overall musicality score and recommending an appropriate musical collaborator. This makes it possible to recommend music content according to the viewer's emotional state and to provide personalized recommendations for the most suitable musical collaborator for the user.

[0406] "Visual and auditory data" refers to video and audio information recorded by the user, and serves as basic data for analyzing movements and musicality related to musical instrument performance.

[0407] "Analysis means" refers to a device or software for extracting movements and characteristics related to musical instrument performance from visual and auditory data and analyzing that information.

[0408] "Acoustic analysis means" refers to a device or software for evaluating sound characteristics from visual and auditory data and determining musicality.

[0409] "Emotional information" refers to data that indicates the emotional state expressed through the user's facial expressions and actions, obtained from visual and auditory data.

[0410] The "Comprehensive Musicality Score" is a numerical value or index that represents a comprehensive evaluation that integrates the results of visual and auditory data analysis and includes musical characteristics and emotional information.

[0411] A "music collaborator" is a partner who collaborates with a user on musical activities, recommended based on emotional information and musicality.

[0412] This system consists of client-side smart devices and server-side data processing facilities. The smart devices are responsible for recording the user's performance as visual and auditory data. Specifically, smartphones and tablets are used, and their cameras and microphones are used to acquire video and audio. The acquired data is transmitted to the server via the internet.

[0413] The server performs various processes using the received visual and auditory data. First, it analyzes the visual data using image processing libraries such as OpenCV to detect playing movements and skeletal movements. This information is used to extract the technical characteristics of the instrument performance. Meanwhile, it utilizes acoustic analysis libraries such as Librosa to analyze the musicality of the acoustic data. This allows for the evaluation of various musical characteristics such as the tempo, rhythm, and melody line of the piece.

[0414] Furthermore, using emotion analysis engines such as Hume AI, the system analyzes the user's facial expressions from visual data to obtain emotional information. This emotional information reflects the user's emotional state during performance and deeply influences the overall musicality score generated by the server. This makes it possible to score not only based on technical skills but also on emotional expressiveness.

[0415] The server matches the user's desired musical style with their overall musicality score and lists suitable musical collaborators. This list is displayed on the user's device, allowing the user to select the collaborator best suited to their musical activities.

[0416] For example, suppose a user plays jazz piano and expresses a sense of exhilaration in their performance. In this case, the system would recommend a musical collaborator who also has an interest in jazz and can play with a similar sense of exhilaration. Another example of a prompt in a generative AI model could be: "Based on the emotional state detected from the user's performance video, recommend jazz content that will give viewers a sense of exhilaration."

[0417] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0418] Step 1: The user records their performance using a smart device and saves it as visual and auditory data. After recording, the device sends this data to a server via the internet. In this step, the video and audio of the performance are inputs, and the data reaches the server as output.

[0419] Step 2: The server analyzes the received visual data using the OpenCV library. This involves detecting the performer's skeletal structure and movements from the video data and obtaining data to evaluate their performance technique. The input is visual data, and the output is the analyzed motion information.

[0420] Step 3: The server analyzes the auditory data using the Librosa library. It extracts musical features, such as tempo and pitch, from the audio data and generates data to evaluate musicality based on these features. The input is auditory data, and the output is the analyzed musical features.

[0421] Step 4: The server uses an emotion analysis engine, such as Hume AI, to analyze visual data and obtain information about the user's emotions. Specifically, it identifies the emotional state during performance from facial expressions and body movements. The input is visual data, and the output is data indicating the emotional state.

[0422] Step 5: The server integrates the previously acquired behavioral information, musical characteristics, and emotional information to calculate an overall musicality score. This score reflects the user's musical skills and emotional expression abilities. The input is the data obtained in Steps 2 through 4, and the output is the overall musicality score.

[0423] Step 6: The server compares the user's desired musicality with the overall musicality score and generates a list to recommend the most suitable musical collaborators. The inputs are the overall musicality score and the user's desired musicality, and the output is a list of collaborators. This list is provided to the user's terminal, enabling matching with appropriate collaborators.

[0424] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0425] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0426] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0427] [Third Embodiment]

[0428] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0429] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0430] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0431] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0432] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0433] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0434] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0435] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0436] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0437] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0438] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0439] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0440] This invention is a system in which users upload videos of their performances to a platform, and an AI model is used to evaluate their performance skills and musical characteristics, and then recommend the most suitable music group members. The following describes an embodiment of this system.

[0441] User actions:

[0442] Users record their performances in video format using devices such as smartphones or PCs. They then upload the recorded videos to the system's dedicated platform. The videos include the performer's movements and audio information.

[0443] Server analysis processing:

[0444] The server first saves the uploaded video to storage. Next, it runs a video analysis AI model to analyze the saved video. The video analysis detects the user's movements and identifies skeletal movements and fingerings. The acoustic analysis module evaluates musical elements based on the audio extracted from the video. This includes variability in tempo, pitch, and volume.

[0445] Score calculation and matching:

[0446] The server integrates the results of motion and acoustic analysis to calculate an overall musicality score. This score quantifies the user's performance skills and musical expressiveness. Next, it compares the calculated score with the user's registered desired musicality. The server then filters the data to recommend suitable music group members to the user and generates a list.

[0447] Specific example:

[0448] For example, suppose a user uploads a video of themselves playing in a "jazz" style. The server analyzes the user's fingering and unique sonic nuances from this video and evaluates their overall musicality score as "Jazz Elegance." The system then selects members who are predicted to be a good match for a user with a "Jazz Elegance" score and presents them to the user as a list.

[0449] In this way, the system of the present invention can effectively match members based on the user's playing skills and musical preferences.

[0450] The following describes the processing flow.

[0451] Step 1:

[0452] Users record their own performances using devices such as smartphones or PCs, and save them as video files.

[0453] Step 2:

[0454] The device selects the video recorded by the user and opens the platform's dedicated upload screen. The user presses the upload button to send the video to the platform.

[0455] Step 3:

[0456] The server receives the uploaded video data and saves it to storage. This prepares it for subsequent analysis.

[0457] Step 4:

[0458] The server starts an AI video analysis model to analyze the stored video. This analysis detects the user's skeletal structure and finger movements from the video.

[0459] Step 5:

[0460] The server extracts audio data from the video and uses an acoustic analysis module to analyze tempo, pitch, and volume changes. This quantifies the musical elements.

[0461] Step 6:

[0462] The server integrates the results of video and audio analysis to calculate an overall musicality score. This score reflects the user's performance skills and musical expressiveness.

[0463] Step 7:

[0464] The server begins matching users with music group members based on their preferred musical style.

[0465] Step 8:

[0466] The server compares the calculated overall musicality score with the user's desired musicality and filters out the most suitable candidates for music group members.

[0467] Step 9:

[0468] The server generates a list of potential music group members based on the filtering results and sends it to the user's terminal.

[0469] Step 10:

[0470] The device displays a list of potential members to the user. The user then decides whether to contact members they are interested in.

[0471] (Example 1)

[0472] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0473] In musical activities, finding suitable musical collaborators who match one's performance skills and musical characteristics is difficult. Therefore, there is a need for a method to efficiently evaluate a performer's abilities and musical characteristics and recommend the most suitable musical collaborators.

[0474] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0475] In this invention, the server includes means for receiving and storing multimedia data recorded by a user's performance, processing means for analyzing the performance actions from the multimedia data, and sound processing means for analyzing acoustic information from the multimedia data and evaluating musical characteristics. This makes it possible to accurately evaluate the performer's skill and musical characteristics and quickly recommend the most suitable musical collaborator to the user.

[0476] A "user" refers to an individual or group that uses the system to record their performances and receive evaluations and recommendations for their musicality.

[0477] "Multimedia data" refers to complex information data that includes audio and video, and specifically data that includes the actions of performers and the sounds of their performances.

[0478] "Performance actions" refer to the physical actions involved in playing a musical instrument, including body movements and finger movements.

[0479] "Processing means" refers to methods and techniques for analyzing and interpreting data in order to achieve a specific purpose.

[0480] "Acoustic information" refers to data that includes sound characteristics and elements related to music, such as tempo, pitch, and volume.

[0481] "Musical characteristics" refer to elements that indicate distinctive style and expressiveness in music, and serve as criteria for evaluating musicality.

[0482] "Acoustic processing means" refers to methods and equipment for effectively analyzing and evaluating acoustic information.

[0483] A "musical collaborator" refers to a partner or group member best suited to the user's performance skills and musical characteristics.

[0484] This invention is a system that allows users to record their own performances, evaluate their musical characteristics, and recommend appropriate musical collaborators. Specific embodiments of the invention are described below.

[0485] Users record their performances in video format using devices such as smartphones or PCs. Typically, high-resolution cameras and high-quality microphones are used for this purpose. The recorded video contains information about the performer's movements and audio, and this data is treated as multimedia data.

[0486] Users upload captured multimedia data to a dedicated platform on a server via an internet connection. The server acts as temporary storage for this multimedia data. Encryption technology is used on this platform to ensure data security.

[0487] The server uses a video analysis AI model to analyze the user's playing movements. This analysis identifies skeletal movements and finger placement, and recognizes movement patterns. The server also uses an acoustic analysis module as an acoustic processing tool to analyze acoustic information such as tempo, pitch, and volume fluctuations.

[0488] Based on the analysis, the server evaluates the user's musical characteristics and quantifies them as a musicality index. This musicality index is then compared to the user's desired musical characteristics, and a list of recommended musical collaborators is generated.

[0489] For example, if a user uploads a video of themselves playing a jazz-style instrument, the server analyzes the fingering and nuances of the sound and evaluates the musicality as "jazz elegance." Based on this, the server lists and recommends to the user musical collaborators who are a good match for the user and who share the "jazz elegance" characteristic.

[0490] An example of a prompt message is, "Evaluate the musicality of the jazz-style performance video and recommend suitable music group members." In this way, the system of the present invention can recommend effective musical collaborators based on the user's performance skills and musical characteristics.

[0491] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0492] Step 1:

[0493] Users record their performances in video format using their smartphones or PCs and upload the video data to the system's dedicated platform. The input is the performance video data, and the output is the video data stored on the platform. Specifically, the user presses the start recording button to begin playing, and then clicks the upload button to send the data after finishing.

[0494] Step 2:

[0495] The server stores the uploaded video data in temporary storage. The input is the video data received from the user, and the output is the file path in the storage. At this stage, the server checks the format and integrity of the video data and stores it in a secure environment. Specifically, it verifies the file format and writes the data to the storage.

[0496] Step 3:

[0497] The server uses a video analysis AI model to analyze performance movements from video data. The input is video data stored in storage, and the output is the analyzed movement data. Specifically, the AI ​​model recognizes the performer's skeletal structure and finger movements frame by frame and extracts movement patterns as digital data.

[0498] Step 4:

[0499] The server uses acoustic processing capabilities to analyze acoustic information from video data. The input is audio data stored in storage, and the output is acoustic characteristic data resulting from the analysis. Specifically, the acoustic processing module acquires information on tempo, pitch, and volume fluctuations and processes the data to evaluate musical characteristics.

[0500] Step 5:

[0501] The server integrates the analysis results of performance actions and acoustic characteristics to calculate an overall musicality index. The input consists of performance data and acoustic characteristic data, and the output is a musicality index score. Specifically, it compares each analysis result, weights them based on specific criteria, and generates an overall score.

[0502] Step 6:

[0503] The server generates a list of recommended musicians based on the calculated musicality index. The input is the musicality index score, and the output is a list of recommended musicians. Specifically, the server's filtering algorithm creates a list by matching it with the user's desired characteristics, and this list is displayed on the user's terminal.

[0504] (Application Example 1)

[0505] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0506] Individual musicians often find it difficult to accurately assess their own performance abilities and musical preferences. Furthermore, the lack of objective criteria for finding suitable musical group members results in a time-consuming and laborious process. This invention aims to provide a system that automatically analyzes a user's performance abilities and musical characteristics and quickly recommends suitable collaborators based on the results.

[0507] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0508] In this invention, the server includes a device for receiving and storing video data recorded by a user, a device for analyzing movements related to the performance from the video data, and a device for analyzing sound from the video data and evaluating musicality. This makes it possible to efficiently analyze the user's performance ability and musical characteristics and recommend appropriate collaborators.

[0509] A "user" is the entity that records video and image data and uploads it to the system.

[0510] "Performing" is the act of expressing music using a musical instrument.

[0511] "Video data" refers to data in video format that records a performance.

[0512] A "storage device" is a mechanism for storing received video and image data.

[0513] An "analysis device" is a technology for detecting and interpreting movements related to a performance from video data.

[0514] A "device for evaluating musicality" is a means of analyzing sound and quantifying the technical and expressive characteristics of a performance.

[0515] "Musicality evaluation" is an indicator of overall performance ability obtained through video data and acoustic analysis.

[0516] A "collaborator" is a member who can work together on musical activities.

[0517] A "recommended device" is a system that suggests appropriate collaborators based on the user's musicality evaluation.

[0518] The "presentation device" is an interface for informing the user of recommended collaborators.

[0519] The system for implementing this invention consists of a user recording their performance using a portable information terminal such as a smartphone and uploading the video data to a server. The server first stores the received video data. Then, it uses a cloud-based analysis service to analyze the performer's movements and music data in the video. This analysis includes a video analysis device and specific software modules for evaluating musicality.

[0520] The server uses video analysis software such as Google Cloud Video Intelligence and AWS Rekognition to analyze the performer's technique and physical structure from video data. Furthermore, it extracts musical characteristics from audio data using Google Cloud Speech-to-Text and Amazon Transcribe to evaluate musicality. These results are integrated to generate a comprehensive musicality evaluation.

[0521] Subsequently, the server uses a generative AI model to extract the most suitable collaborators based on musicality evaluations and generates a user-specific recommendation list. This process includes a scoring algorithm using a programming language such as Python. On the user's device, an interface is built to receive and display the recommendation list from the server.

[0522] For example, if a user uploads a video of themselves playing classical piano, the server will analyze it, evaluate the user's playing technique according to classical music standards, and recommend other performers with similar musicality. This analysis can be performed using a prompt message such as, "This video contains a performance of classical music. Please analyze the fingering accuracy and musical expressiveness."

[0523] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0524] Step 1:

[0525] Users record their performances as videos on devices such as smartphones and upload them to the server using the system's application. The input is the user's performance video file, and the output is the video data stored on the server.

[0526] Step 2:

[0527] The server passes the received video data to video analysis software (e.g., Google Cloud Video Intelligence) to analyze the performer's technique and physical structure. In this step, the input is video data, and the output is analysis information (e.g., finger movements and rhythm patterns).

[0528] Step 3:

[0529] The server extracts audio data and performs acoustic analysis. The audio data is then fed into Google Cloud Speech-to-Text or Amazon Transcribe to extract musical features (e.g., tempo, pitch). The input is audio data, and the output is musical feature data.

[0530] Step 4:

[0531] The server integrates data obtained from video analysis software and audio analysis software to generate a comprehensive musicality evaluation. It uses video analysis information and musical characteristic data as input, and outputs a musicality evaluation score.

[0532] Step 5:

[0533] Using a generative AI model, the server lists suitable collaborators based on musicality evaluation scores. In this step, the input is the musicality evaluation score, and the output is a recommendation list. The prompt message can be "This video contains a performance of classical music. Please analyze the fingering accuracy and musical expressiveness."

[0534] Step 6:

[0535] The server sends the generated recommendation list to the user's terminal, which then launches an interface to display the list. The input is the recommendation list, and the output is a list of collaborators displayed on the terminal's screen.

[0536] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0537] This invention combines an emotion engine with a system that comprehensively analyzes data obtained from users' performance videos and recommends appropriate members for a music group, thereby achieving more personalized matching that takes into account the user's emotional state. The embodiments of this system are described in detail below.

[0538] User actions:

[0539] Users record their performances using devices such as smartphones or PCs. The recordings include both video and audio of the performance, and users upload these video files to the platform.

[0540] Server processing:

[0541] The server first saves the received video data to storage. For the video, an AI model for video analysis is used to perform skeletal detection and finger movement analysis, and for the audio, an acoustic analysis module is used to evaluate the characteristics of the music. In this process, an emotion engine is also used to analyze the user's facial expressions and movements from the video and obtain emotion data.

[0542] Score calculation and adjustment:

[0543] The server integrates video analysis, acoustic analysis, and emotional data to calculate an overall musicality score. Based on the emotional data obtained by the emotion engine, the score can be adjusted to take into account the user's emotional state.

[0544] Matching process:

[0545] The server uses the calculated overall musicality score to match the user with their desired musical style. Based on emotional data, it further refines the filtering process to generate a list of suitable music group members for the user.

[0546] Specific example:

[0547] For example, if a user uploads a video of themselves playing the blues, the emotion engine might recognize that the user experienced high levels of satisfaction during their performance. The server would then consider this positive emotion data and list other members who share similar emotional expressiveness with other members specializing in the blues. This emotional data-driven matching process can foster richer musical interaction between the user and existing members.

[0548] This system enables personalized music group matching that takes into account the user's emotional state, in addition to traditional technical and musical evaluations.

[0549] The following describes the processing flow.

[0550] Step 1:

[0551] Users record their performances as videos on their smartphones or PCs. These videos should capture the performer's facial expressions and movements.

[0552] Step 2:

[0553] The device sends the recorded video file to the server by opening the platform's dedicated upload screen, selecting the file, and pressing the upload button.

[0554] Step 3:

[0555] The server receives the uploaded video data and saves it to storage. This storage enables subsequent analysis processes.

[0556] Step 4:

[0557] The server inputs the saved videos into an AI video analysis model to analyze the performer's skeletal structure and finger movements. This allows for the quantification of the characteristics of the performance technique.

[0558] Step 5:

[0559] The server passes the audio data extracted from the video to an audio analysis module, which analyzes the tempo, pitch, and volume changes of the music. The musical characteristics are then quantified.

[0560] Step 6:

[0561] The server activates an emotion engine and analyzes the user's facial expressions and gestures in the video. This allows the emotional state during the performance to be quantified.

[0562] Step 7:

[0563] The server integrates the results of video analysis, audio analysis, and emotion engine analysis to calculate an overall musicality score. The score is then adjusted based on the recognized emotion data.

[0564] Step 8:

[0565] The server compares the user's registered musical preferences with an overall musicality score to search for suitable music group members. It further narrows down the candidates by utilizing emotional data.

[0566] Step 9:

[0567] The server generates a list of optimized music group member candidates and sends it to the user's terminal.

[0568] Step 10:

[0569] The device displays a list of potential members to the user. The user then decides whether to begin interacting with members from that list who interest them.

[0570] (Example 2)

[0571] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0572] In musical activities, for performers to collaborate with other members, not only technical ability but also musicality and emotional compatibility are important. However, conventional systems do not take such emotional compatibility into consideration when making recommendations, which can lead to inappropriate matches. Therefore, in order to improve users' satisfaction with their musical activities, there is a need for recommendations of music group members that take emotional aspects into account.

[0573] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0574] In this invention, the server includes means for receiving and storing video footage recorded by the user, means for analyzing actions related to musical instrument performance from the video footage, and acoustic analysis means for analyzing sound from the video footage and evaluating musicality. This makes it possible to adjust the overall musicality evaluation based on emotional data and recommend members of a musical group that take into account the user's emotional state.

[0575] A "user" is the entity that uses the system to record a performance and uploads that data.

[0576] "Video and audio materials" refers to video files recorded by users during performances, and includes both video and audio.

[0577] "Skeleton" refers to a data structure that shows the location of the main bones and joints of a performer's body within video footage.

[0578] "Emotional data" refers to information indicating the emotional state extracted from the user's facial expressions and movements, and is used to evaluate emotional expression during performance.

[0579] "Comprehensive Musical Quality Evaluation" is an evaluation index that shows musical quality by integrating video analysis results, sound analysis results, and emotional data.

[0580] "Music group members" refers to other members who are encouraged to collaborate with the user in musical activities.

[0581] The "list" is a compilation of candidate members for a music group, selected based on the user's playing style and emotional state.

[0582] A "server" is a computer device that serves as the core of the entire system, responsible for data storage, analysis, and the generation of recommendation results.

[0583] This system uses video recordings of user performances to appropriately match users with other members of musical groups. Users record their performances using devices such as smartphones or PCs and upload the video data to the platform. The recorded videos include both video and audio of the performance. The system aims to recommend the most suitable members for a musical group based on the user's performance style and emotions.

[0584] The server is equipped with an AI model that analyzes video footage and evaluates the user's movements and musical characteristics during performance. First, the server saves the uploaded video data to its storage. Next, it uses the video analysis AI model to perform skeletal detection and finger movement analysis of the user. This analysis allows for a precise understanding of the user's movements related to playing the instrument.

[0585] The acoustic analysis module analyzes audio data and evaluates the tempo, rhythm, and musicality of the music. Based on these analysis results, the system identifies the user's musical genre and playing style.

[0586] Furthermore, an emotion engine is used to acquire emotional data from the user's facial expressions and actions within the video. This makes it possible to adjust the overall musicality evaluation while considering the user's emotional state. The emotional data reflects the user's emotional expression during performance and is used to filter out appropriate musical components.

[0587] As a concrete example, consider a scenario where a user uploads a video of themselves playing "jazz," and the emotion engine recognizes high levels of excitement and pride in the user's performance. Based on this information, the server creates a list of members with similar emotional expressiveness specializing in jazz performance. This process facilitates rich musical interaction based on emotional data.

[0588] An example of a prompt message is: "We will analyze the subject's face and movements from this user's performance video, identify their emotional state during the performance, and use that data to adjust the overall musicality evaluation. Then, based on the adjusted evaluation, we will recommend suitable music group members."

[0589] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0590] Step 1:

[0591] Users record their performances using devices such as smartphones or PCs and upload the video data to the system's platform. The input here is a video file containing both video and audio. This file forms the basis for subsequent analysis steps.

[0592] Step 2:

[0593] The server receives video files uploaded by users and saves them to storage. During this saving process, the server checks the format of the video files and converts them to a format suitable for analysis if necessary. This conversion ensures that the server has access to the data needed for analysis.

[0594] Step 3:

[0595] The server activates a video analysis AI model to analyze the video portion of the footage. Here, a skeletal detection algorithm is used to identify the movements of the performer's body and fingers, and fingering data is extracted. The input is video data, and the output is quantitative data indicating the performer's movements.

[0596] Step 4:

[0597] The server uses an acoustic analysis module to analyze the audio portion of a video and evaluate its musical characteristics, such as tempo, rhythm, and genre. The input for this process is audio data obtained from the video, and the output is data that describes the musical characteristics in detail.

[0598] Step 5:

[0599] The server uses an emotion engine to acquire emotional data based on the user's facial expressions and actions in the video. The input is video data, and the output is qualitative data indicating the user's emotional state. This information is used to understand emotional expression during performance.

[0600] Step 6:

[0601] The server integrates the previously obtained video analysis results, acoustic analysis results, and emotion data to calculate an overall musicality rating. In this integration process, individual data points are converted into evaluation scales that indicate the user's musical skills and emotional expressiveness. The output is an overall musicality score.

[0602] Step 7:

[0603] The server performs a process of recommending the most suitable members of a musical group to the user based on an overall musicality evaluation. It also considers emotional compatibility based on emotional data, filters the recommendations, generates a final recommendation list, and sends it to the user's terminal. The input is the overall musicality evaluation, and the output is a list of recommended musical members.

[0604] (Application Example 2)

[0605] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0606] Traditional music content recommendation systems primarily rely on recommendations based on musical skill and preferences, making it difficult to provide personalized recommendations that reflect the listener's emotional state. Furthermore, collaborator recommendations, designed to help users maximize their musical potential, also lack emotional considerations for personalization. Therefore, there are limitations to providing the optimal music experience for listeners, and a challenge remains in enhancing users' emotional satisfaction with their musical activities.

[0607] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0608] In this invention, the server includes means for receiving and storing visual and auditory data recorded by the user; analysis means for analyzing movements related to musical instrument performance from the visual and auditory data; acoustic analysis means for analyzing sound from the visual and auditory data and evaluating musicality; means for analyzing emotional state from the visual and auditory data and acquiring emotional information; means for integrating the results of the video analysis, acoustic analysis, and emotional information to calculate an overall musicality score; and means for comparing the user's desired musicality with the overall musicality score and recommending an appropriate musical collaborator. This makes it possible to recommend music content according to the viewer's emotional state and to provide personalized recommendations for the most suitable musical collaborator for the user.

[0609] "Visual and auditory data" refers to video and audio information recorded by the user, and serves as basic data for analyzing movements and musicality related to musical instrument performance.

[0610] "Analysis means" refers to a device or software for extracting movements and characteristics related to musical instrument performance from visual and auditory data and analyzing that information.

[0611] "Acoustic analysis means" refers to a device or software for evaluating sound characteristics from visual and auditory data and determining musicality.

[0612] "Emotional information" refers to data that indicates the emotional state expressed through the user's facial expressions and actions, obtained from visual and auditory data.

[0613] The "Comprehensive Musicality Score" is a numerical value or index that represents a comprehensive evaluation that integrates the results of visual and auditory data analysis and includes musical characteristics and emotional information.

[0614] A "music collaborator" is a partner who collaborates with a user on musical activities, recommended based on emotional information and musicality.

[0615] This system consists of client-side smart devices and server-side data processing facilities. The smart devices are responsible for recording the user's performance as visual and auditory data. Specifically, smartphones and tablets are used, and their cameras and microphones are used to acquire video and audio. The acquired data is transmitted to the server via the internet.

[0616] The server performs various processes using the received visual and auditory data. First, it analyzes the visual data using image processing libraries such as OpenCV to detect playing movements and skeletal movements. This information is used to extract the technical characteristics of the instrument performance. Meanwhile, it utilizes acoustic analysis libraries such as Librosa to analyze the musicality of the acoustic data. This allows for the evaluation of various musical characteristics such as the tempo, rhythm, and melody line of the piece.

[0617] Furthermore, using emotion analysis engines such as Hume AI, the system analyzes the user's facial expressions from visual data to obtain emotional information. This emotional information reflects the user's emotional state during performance and deeply influences the overall musicality score generated by the server. This makes it possible to score not only based on technical skills but also on emotional expressiveness.

[0618] The server matches the user's desired musical style with their overall musicality score and lists suitable musical collaborators. This list is displayed on the user's device, allowing the user to select the collaborator best suited to their musical activities.

[0619] For example, suppose a user plays jazz piano and expresses a sense of exhilaration in their performance. In this case, the system would recommend a musical collaborator who also has an interest in jazz and can play with a similar sense of exhilaration. Another example of a prompt in a generative AI model could be: "Based on the emotional state detected from the user's performance video, recommend jazz content that will give viewers a sense of exhilaration."

[0620] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0621] Step 1: The user records their performance using a smart device and saves it as visual and auditory data. After recording, the device sends this data to a server via the internet. In this step, the video and audio of the performance are inputs, and the data reaches the server as output.

[0622] Step 2: The server analyzes the received visual data using the OpenCV library. This involves detecting the performer's skeletal structure and movements from the video data and obtaining data to evaluate their performance technique. The input is visual data, and the output is the analyzed motion information.

[0623] Step 3: The server analyzes the auditory data using the Librosa library. It extracts musical features, such as tempo and pitch, from the audio data and generates data to evaluate musicality based on these features. The input is auditory data, and the output is the analyzed musical features.

[0624] Step 4: The server uses an emotion analysis engine, such as Hume AI, to analyze visual data and obtain information about the user's emotions. Specifically, it identifies the emotional state during performance from facial expressions and body movements. The input is visual data, and the output is data indicating the emotional state.

[0625] Step 5: The server integrates the previously acquired behavioral information, musical characteristics, and emotional information to calculate an overall musicality score. This score reflects the user's musical skills and emotional expression abilities. The input is the data obtained in Steps 2 through 4, and the output is the overall musicality score.

[0626] Step 6: The server compares the user's desired musicality with the overall musicality score and generates a list to recommend the most suitable musical collaborators. The inputs are the overall musicality score and the user's desired musicality, and the output is a list of collaborators. This list is provided to the user's terminal, enabling matching with appropriate collaborators.

[0627] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0628] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0629] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0630] [Fourth Embodiment]

[0631] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0632] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0633] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0634] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0635] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0636] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0637] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0638] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0639] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0640] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0641] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0642] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0643] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0644] This invention is a system in which users upload videos of their performances to a platform, and an AI model is used to evaluate their performance skills and musical characteristics, and then recommend the most suitable music group members. The following describes an embodiment of this system.

[0645] User actions:

[0646] Users record their performances in video format using devices such as smartphones or PCs. They then upload the recorded videos to the system's dedicated platform. The videos include the performer's movements and audio information.

[0647] Server analysis processing:

[0648] The server first saves the uploaded video to storage. Next, it runs a video analysis AI model to analyze the saved video. The video analysis detects the user's movements and identifies skeletal movements and fingerings. The acoustic analysis module evaluates musical elements based on the audio extracted from the video. This includes variability in tempo, pitch, and volume.

[0649] Score calculation and matching:

[0650] The server integrates the results of motion and acoustic analysis to calculate an overall musicality score. This score quantifies the user's performance skills and musical expressiveness. Next, it compares the calculated score with the user's registered desired musicality. The server then filters the data to recommend suitable music group members to the user and generates a list.

[0651] Specific example:

[0652] For example, suppose a user uploads a video of themselves playing in a "jazz" style. The server analyzes the user's fingering and unique sonic nuances from this video and evaluates their overall musicality score as "Jazz Elegance." The system then selects members who are predicted to be a good match for a user with a "Jazz Elegance" score and presents them to the user as a list.

[0653] In this way, the system of the present invention can effectively match members based on the user's playing skills and musical preferences.

[0654] The following describes the processing flow.

[0655] Step 1:

[0656] Users record their own performances using devices such as smartphones or PCs, and save them as video files.

[0657] Step 2:

[0658] The device selects the video recorded by the user and opens the platform's dedicated upload screen. The user presses the upload button to send the video to the platform.

[0659] Step 3:

[0660] The server receives the uploaded video data and saves it to storage. This prepares it for subsequent analysis.

[0661] Step 4:

[0662] The server starts an AI video analysis model to analyze the stored video. This analysis detects the user's skeletal structure and finger movements from the video.

[0663] Step 5:

[0664] The server extracts audio data from the video and uses an acoustic analysis module to analyze tempo, pitch, and volume changes. This quantifies the musical elements.

[0665] Step 6:

[0666] The server integrates the results of video and audio analysis to calculate an overall musicality score. This score reflects the user's performance skills and musical expressiveness.

[0667] Step 7:

[0668] The server begins matching users with music group members based on their preferred musical style.

[0669] Step 8:

[0670] The server compares the calculated overall musicality score with the user's desired musicality and filters out the most suitable candidates for music group members.

[0671] Step 9:

[0672] The server generates a list of potential music group members based on the filtering results and sends it to the user's terminal.

[0673] Step 10:

[0674] The device displays a list of potential members to the user. The user then decides whether to contact members they are interested in.

[0675] (Example 1)

[0676] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0677] In musical activities, finding suitable musical collaborators who match one's performance skills and musical characteristics is difficult. Therefore, there is a need for a method to efficiently evaluate a performer's abilities and musical characteristics and recommend the most suitable musical collaborators.

[0678] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0679] In this invention, the server includes means for receiving and storing multimedia data recorded by a user's performance, processing means for analyzing the performance actions from the multimedia data, and sound processing means for analyzing acoustic information from the multimedia data and evaluating musical characteristics. This makes it possible to accurately evaluate the performer's skill and musical characteristics and quickly recommend the most suitable musical collaborator to the user.

[0680] A "user" refers to an individual or group that uses the system to record their performances and receive evaluations and recommendations for their musicality.

[0681] "Multimedia data" refers to complex information data that includes audio and video, and specifically data that includes the actions of performers and the sounds of their performances.

[0682] "Performance actions" refer to the physical actions involved in playing a musical instrument, including body movements and finger movements.

[0683] "Processing means" refers to methods and techniques for analyzing and interpreting data in order to achieve a specific purpose.

[0684] "Acoustic information" refers to data that includes sound characteristics and elements related to music, such as tempo, pitch, and volume.

[0685] "Musical characteristics" refer to elements that indicate distinctive style and expressiveness in music, and serve as criteria for evaluating musicality.

[0686] "Acoustic processing means" refers to methods and equipment for effectively analyzing and evaluating acoustic information.

[0687] A "musical collaborator" refers to a partner or group member best suited to the user's performance skills and musical characteristics.

[0688] This invention is a system that allows users to record their own performances, evaluate their musical characteristics, and recommend appropriate musical collaborators. Specific embodiments of the invention are described below.

[0689] Users record their performances in video format using devices such as smartphones or PCs. Typically, high-resolution cameras and high-quality microphones are used for this purpose. The recorded video contains information about the performer's movements and audio, and this data is treated as multimedia data.

[0690] Users upload captured multimedia data to a dedicated platform on a server via an internet connection. The server acts as temporary storage for this multimedia data. Encryption technology is used on this platform to ensure data security.

[0691] The server uses a video analysis AI model to analyze the user's playing movements. This analysis identifies skeletal movements and finger placement, and recognizes movement patterns. The server also uses an acoustic analysis module as an acoustic processing tool to analyze acoustic information such as tempo, pitch, and volume fluctuations.

[0692] Based on the analysis, the server evaluates the user's musical characteristics and quantifies them as a musicality index. This musicality index is then compared to the user's desired musical characteristics, and a list of recommended musical collaborators is generated.

[0693] For example, if a user uploads a video of themselves playing a jazz-style instrument, the server analyzes the fingering and nuances of the sound and evaluates the musicality as "jazz elegance." Based on this, the server lists and recommends to the user musical collaborators who are a good match for the user and who share the "jazz elegance" characteristic.

[0694] An example of a prompt message is, "Evaluate the musicality of the jazz-style performance video and recommend suitable music group members." In this way, the system of the present invention can recommend effective musical collaborators based on the user's performance skills and musical characteristics.

[0695] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0696] Step 1:

[0697] Users record their performances in video format using their smartphones or PCs and upload the video data to the system's dedicated platform. The input is the performance video data, and the output is the video data stored on the platform. Specifically, the user presses the start recording button to begin playing, and then clicks the upload button to send the data after finishing.

[0698] Step 2:

[0699] The server stores the uploaded video data in temporary storage. The input is the video data received from the user, and the output is the file path in the storage. At this stage, the server checks the format and integrity of the video data and stores it in a secure environment. Specifically, it verifies the file format and writes the data to the storage.

[0700] Step 3:

[0701] The server uses a video analysis AI model to analyze performance movements from video data. The input is video data stored in storage, and the output is the analyzed movement data. Specifically, the AI ​​model recognizes the performer's skeletal structure and finger movements frame by frame and extracts movement patterns as digital data.

[0702] Step 4:

[0703] The server uses acoustic processing capabilities to analyze acoustic information from video data. The input is audio data stored in storage, and the output is acoustic characteristic data resulting from the analysis. Specifically, the acoustic processing module acquires information on tempo, pitch, and volume fluctuations and processes the data to evaluate musical characteristics.

[0704] Step 5:

[0705] The server integrates the analysis results of performance actions and acoustic characteristics to calculate an overall musicality index. The input consists of performance data and acoustic characteristic data, and the output is a musicality index score. Specifically, it compares each analysis result, weights them based on specific criteria, and generates an overall score.

[0706] Step 6:

[0707] The server generates a list of recommended musicians based on the calculated musicality index. The input is the musicality index score, and the output is a list of recommended musicians. Specifically, the server's filtering algorithm creates a list by matching it with the user's desired characteristics, and this list is displayed on the user's terminal.

[0708] (Application Example 1)

[0709] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0710] Individual musicians often find it difficult to accurately assess their own performance abilities and musical preferences. Furthermore, the lack of objective criteria for finding suitable musical group members results in a time-consuming and laborious process. This invention aims to provide a system that automatically analyzes a user's performance abilities and musical characteristics and quickly recommends suitable collaborators based on the results.

[0711] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0712] In this invention, the server includes a device for receiving and storing video data recorded by a user, a device for analyzing movements related to the performance from the video data, and a device for analyzing sound from the video data and evaluating musicality. This makes it possible to efficiently analyze the user's performance ability and musical characteristics and recommend appropriate collaborators.

[0713] A "user" is the entity that records video and image data and uploads it to the system.

[0714] "Performing" is the act of expressing music using a musical instrument.

[0715] "Video data" refers to data in video format that records a performance.

[0716] A "storage device" is a mechanism for storing received video and image data.

[0717] An "analysis device" is a technology for detecting and interpreting movements related to a performance from video data.

[0718] A "device for evaluating musicality" is a means of analyzing sound and quantifying the technical and expressive characteristics of a performance.

[0719] "Musicality evaluation" is an indicator of overall performance ability obtained through video data and acoustic analysis.

[0720] A "collaborator" is a member who can work together on musical activities.

[0721] A "recommended device" is a system that suggests appropriate collaborators based on the user's musicality evaluation.

[0722] The "presentation device" is an interface for informing the user of recommended collaborators.

[0723] The system for implementing this invention consists of a user recording their performance using a portable information terminal such as a smartphone and uploading the video data to a server. The server first stores the received video data. Then, it uses a cloud-based analysis service to analyze the performer's movements and music data in the video. This analysis includes a video analysis device and specific software modules for evaluating musicality.

[0724] The server uses video analysis software such as Google Cloud Video Intelligence and AWS Rekognition to analyze the performer's technique and physical structure from video data. Furthermore, it extracts musical characteristics from audio data using Google Cloud Speech-to-Text and Amazon Transcribe to evaluate musicality. These results are integrated to generate a comprehensive musicality evaluation.

[0725] Subsequently, the server uses a generative AI model to extract the most suitable collaborators based on musicality evaluations and generates a user-specific recommendation list. This process includes a scoring algorithm using a programming language such as Python. On the user's device, an interface is built to receive and display the recommendation list from the server.

[0726] For example, if a user uploads a video of themselves playing classical piano, the server will analyze it, evaluate the user's playing technique according to classical music standards, and recommend other performers with similar musicality. This analysis can be performed using a prompt message such as, "This video contains a performance of classical music. Please analyze the fingering accuracy and musical expressiveness."

[0727] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0728] Step 1:

[0729] Users record their performances as videos on devices such as smartphones and upload them to the server using the system's application. The input is the user's performance video file, and the output is the video data stored on the server.

[0730] Step 2:

[0731] The server passes the received video data to video analysis software (e.g., Google Cloud Video Intelligence) to analyze the performer's technique and physical structure. In this step, the input is video data, and the output is analysis information (e.g., finger movements and rhythm patterns).

[0732] Step 3:

[0733] The server extracts audio data and performs acoustic analysis. The audio data is then fed into Google Cloud Speech-to-Text or Amazon Transcribe to extract musical features (e.g., tempo, pitch). The input is audio data, and the output is musical feature data.

[0734] Step 4:

[0735] The server integrates data obtained from video analysis software and audio analysis software to generate a comprehensive musicality evaluation. It uses video analysis information and musical characteristic data as input, and outputs a musicality evaluation score.

[0736] Step 5:

[0737] Using a generative AI model, the server lists suitable collaborators based on musicality evaluation scores. In this step, the input is the musicality evaluation score, and the output is a recommendation list. The prompt message can be "This video contains a performance of classical music. Please analyze the fingering accuracy and musical expressiveness."

[0738] Step 6:

[0739] The server sends the generated recommendation list to the user's terminal, which then launches an interface to display the list. The input is the recommendation list, and the output is a list of collaborators displayed on the terminal's screen.

[0740] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0741] This invention combines an emotion engine with a system that comprehensively analyzes data obtained from users' performance videos and recommends appropriate members for a music group, thereby achieving more personalized matching that takes into account the user's emotional state. The embodiments of this system are described in detail below.

[0742] User actions:

[0743] Users record their performances using devices such as smartphones or PCs. The recordings include both video and audio of the performance, and users upload these video files to the platform.

[0744] Server processing:

[0745] The server first saves the received video data to storage. For the video, an AI model for video analysis is used to perform skeletal detection and finger movement analysis, and for the audio, an acoustic analysis module is used to evaluate the characteristics of the music. In this process, an emotion engine is also used to analyze the user's facial expressions and movements from the video and obtain emotion data.

[0746] Score calculation and adjustment:

[0747] The server integrates video analysis, acoustic analysis, and emotional data to calculate an overall musicality score. Based on the emotional data obtained by the emotion engine, the score can be adjusted to take into account the user's emotional state.

[0748] Matching process:

[0749] The server uses the calculated overall musicality score to match the user with their desired musical style. Based on emotional data, it further refines the filtering process to generate a list of suitable music group members for the user.

[0750] Specific example:

[0751] For example, if a user uploads a video of themselves playing the blues, the emotion engine might recognize that the user experienced high levels of satisfaction during their performance. The server would then consider this positive emotion data and list other members who share similar emotional expressiveness with other members specializing in the blues. This emotional data-driven matching process can foster richer musical interaction between the user and existing members.

[0752] This system enables personalized music group matching that takes into account the user's emotional state, in addition to traditional technical and musical evaluations.

[0753] The following describes the processing flow.

[0754] Step 1:

[0755] Users record their performances as videos on their smartphones or PCs. These videos should capture the performer's facial expressions and movements.

[0756] Step 2:

[0757] The device sends the recorded video file to the server by opening the platform's dedicated upload screen, selecting the file, and pressing the upload button.

[0758] Step 3:

[0759] The server receives the uploaded video data and saves it to storage. This storage enables subsequent analysis processes.

[0760] Step 4:

[0761] The server inputs the saved videos into an AI video analysis model to analyze the performer's skeletal structure and finger movements. This allows for the quantification of the characteristics of the performance technique.

[0762] Step 5:

[0763] The server passes the audio data extracted from the video to an audio analysis module, which analyzes the tempo, pitch, and volume changes of the music. The musical characteristics are then quantified.

[0764] Step 6:

[0765] The server activates an emotion engine and analyzes the user's facial expressions and gestures in the video. This allows the emotional state during the performance to be quantified.

[0766] Step 7:

[0767] The server integrates the results of video analysis, audio analysis, and emotion engine analysis to calculate an overall musicality score. The score is then adjusted based on the recognized emotion data.

[0768] Step 8:

[0769] The server compares the user's registered musical preferences with an overall musicality score to search for suitable music group members. It further narrows down the candidates by utilizing emotional data.

[0770] Step 9:

[0771] The server generates a list of optimized music group member candidates and sends it to the user's terminal.

[0772] Step 10:

[0773] The device displays a list of potential members to the user. The user then decides whether to begin interacting with members from that list who interest them.

[0774] (Example 2)

[0775] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0776] In musical activities, for performers to collaborate with other members, not only technical ability but also musicality and emotional compatibility are important. However, conventional systems do not take such emotional compatibility into consideration when making recommendations, which can lead to inappropriate matches. Therefore, in order to improve users' satisfaction with their musical activities, there is a need for recommendations of music group members that take emotional aspects into account.

[0777] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0778] In this invention, the server includes means for receiving and storing video footage recorded by the user, means for analyzing actions related to musical instrument performance from the video footage, and acoustic analysis means for analyzing sound from the video footage and evaluating musicality. This makes it possible to adjust the overall musicality evaluation based on emotional data and recommend members of a musical group that take into account the user's emotional state.

[0779] A "user" is the entity that uses the system to record a performance and uploads that data.

[0780] "Video and audio materials" refers to video files recorded by users during performances, and includes both video and audio.

[0781] "Skeleton" refers to a data structure that shows the location of the main bones and joints of a performer's body within video footage.

[0782] "Emotional data" refers to information indicating the emotional state extracted from the user's facial expressions and movements, and is used to evaluate emotional expression during performance.

[0783] "Comprehensive Musical Quality Evaluation" is an evaluation index that shows musical quality by integrating video analysis results, sound analysis results, and emotional data.

[0784] "Music group members" refers to other members who are encouraged to collaborate with the user in musical activities.

[0785] The "list" is a compilation of candidate members for a music group, selected based on the user's playing style and emotional state.

[0786] A "server" is a computer device that serves as the core of the entire system, responsible for data storage, analysis, and the generation of recommendation results.

[0787] This system uses video recordings of user performances to appropriately match users with other members of musical groups. Users record their performances using devices such as smartphones or PCs and upload the video data to the platform. The recorded videos include both video and audio of the performance. The system aims to recommend the most suitable members for a musical group based on the user's performance style and emotions.

[0788] The server is equipped with an AI model that analyzes video footage and evaluates the user's movements and musical characteristics during performance. First, the server saves the uploaded video data to its storage. Next, it uses the video analysis AI model to perform skeletal detection and finger movement analysis of the user. This analysis allows for a precise understanding of the user's movements related to playing the instrument.

[0789] The acoustic analysis module analyzes audio data and evaluates the tempo, rhythm, and musicality of the music. Based on these analysis results, the system identifies the user's musical genre and playing style.

[0790] Furthermore, an emotion engine is used to acquire emotional data from the user's facial expressions and actions within the video. This makes it possible to adjust the overall musicality evaluation while considering the user's emotional state. The emotional data reflects the user's emotional expression during performance and is used to filter out appropriate musical components.

[0791] As a concrete example, consider a scenario where a user uploads a video of themselves playing "jazz," and the emotion engine recognizes high levels of excitement and pride in the user's performance. Based on this information, the server creates a list of members with similar emotional expressiveness specializing in jazz performance. This process facilitates rich musical interaction based on emotional data.

[0792] An example of a prompt message is: "We will analyze the subject's face and movements from this user's performance video, identify their emotional state during the performance, and use that data to adjust the overall musicality evaluation. Then, based on the adjusted evaluation, we will recommend suitable music group members."

[0793] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0794] Step 1:

[0795] Users record their performances using devices such as smartphones or PCs and upload the video data to the system's platform. The input here is a video file containing both video and audio. This file forms the basis for subsequent analysis steps.

[0796] Step 2:

[0797] The server receives video files uploaded by users and saves them to storage. During this saving process, the server checks the format of the video files and converts them to a format suitable for analysis if necessary. This conversion ensures that the server has access to the data needed for analysis.

[0798] Step 3:

[0799] The server activates a video analysis AI model to analyze the video portion of the footage. Here, a skeletal detection algorithm is used to identify the movements of the performer's body and fingers, and fingering data is extracted. The input is video data, and the output is quantitative data indicating the performer's movements.

[0800] Step 4:

[0801] The server uses an acoustic analysis module to analyze the audio portion of a video and evaluate its musical characteristics, such as tempo, rhythm, and genre. The input for this process is audio data obtained from the video, and the output is data that describes the musical characteristics in detail.

[0802] Step 5:

[0803] The server uses an emotion engine to acquire emotional data based on the user's facial expressions and actions in the video. The input is video data, and the output is qualitative data indicating the user's emotional state. This information is used to understand emotional expression during performance.

[0804] Step 6:

[0805] The server integrates the previously obtained video analysis results, acoustic analysis results, and emotion data to calculate an overall musicality rating. In this integration process, individual data points are converted into evaluation scales that indicate the user's musical skills and emotional expressiveness. The output is an overall musicality score.

[0806] Step 7:

[0807] The server performs a process of recommending the most suitable members of a musical group to the user based on an overall musicality evaluation. It also considers emotional compatibility based on emotional data, filters the recommendations, generates a final recommendation list, and sends it to the user's terminal. The input is the overall musicality evaluation, and the output is a list of recommended musical members.

[0808] (Application Example 2)

[0809] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0810] Traditional music content recommendation systems primarily rely on recommendations based on musical skill and preferences, making it difficult to provide personalized recommendations that reflect the listener's emotional state. Furthermore, collaborator recommendations, designed to help users maximize their musical potential, also lack emotional considerations for personalization. Therefore, there are limitations to providing the optimal music experience for listeners, and a challenge remains in enhancing users' emotional satisfaction with their musical activities.

[0811] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0812] In this invention, the server includes means for receiving and storing visual and auditory data recorded by the user; analysis means for analyzing movements related to musical instrument performance from the visual and auditory data; acoustic analysis means for analyzing sound from the visual and auditory data and evaluating musicality; means for analyzing emotional state from the visual and auditory data and acquiring emotional information; means for integrating the results of the video analysis, acoustic analysis, and emotional information to calculate an overall musicality score; and means for comparing the user's desired musicality with the overall musicality score and recommending an appropriate musical collaborator. This makes it possible to recommend music content according to the viewer's emotional state and to provide personalized recommendations for the most suitable musical collaborator for the user.

[0813] "Visual and auditory data" refers to video and audio information recorded by the user, and serves as basic data for analyzing movements and musicality related to musical instrument performance.

[0814] "Analysis means" refers to a device or software for extracting movements and characteristics related to musical instrument performance from visual and auditory data and analyzing that information.

[0815] "Acoustic analysis means" refers to a device or software for evaluating sound characteristics from visual and auditory data and determining musicality.

[0816] "Emotional information" refers to data that indicates the emotional state expressed through the user's facial expressions and actions, obtained from visual and auditory data.

[0817] The "Comprehensive Musicality Score" is a numerical value or index that represents a comprehensive evaluation that integrates the results of visual and auditory data analysis and includes musical characteristics and emotional information.

[0818] A "music collaborator" is a partner who collaborates with a user on musical activities, recommended based on emotional information and musicality.

[0819] This system consists of client-side smart devices and server-side data processing facilities. The smart devices are responsible for recording the user's performance as visual and auditory data. Specifically, smartphones and tablets are used, and their cameras and microphones are used to acquire video and audio. The acquired data is transmitted to the server via the internet.

[0820] The server performs various processes using the received visual and auditory data. First, it analyzes the visual data using image processing libraries such as OpenCV to detect playing movements and skeletal movements. This information is used to extract the technical characteristics of the instrument performance. Meanwhile, it utilizes acoustic analysis libraries such as Librosa to analyze the musicality of the acoustic data. This allows for the evaluation of various musical characteristics such as the tempo, rhythm, and melody line of the piece.

[0821] Furthermore, using emotion analysis engines such as Hume AI, the system analyzes the user's facial expressions from visual data to obtain emotional information. This emotional information reflects the user's emotional state during performance and deeply influences the overall musicality score generated by the server. This makes it possible to score not only based on technical skills but also on emotional expressiveness.

[0822] The server matches the user's desired musical style with their overall musicality score and lists suitable musical collaborators. This list is displayed on the user's device, allowing the user to select the collaborator best suited to their musical activities.

[0823] For example, suppose a user plays jazz piano and expresses a sense of exhilaration in their performance. In this case, the system would recommend a musical collaborator who also has an interest in jazz and can play with a similar sense of exhilaration. Another example of a prompt in a generative AI model could be: "Based on the emotional state detected from the user's performance video, recommend jazz content that will give viewers a sense of exhilaration."

[0824] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0825] Step 1: The user records their performance using a smart device and saves it as visual and auditory data. After recording, the device sends this data to a server via the internet. In this step, the video and audio of the performance are inputs, and the data reaches the server as output.

[0826] Step 2: The server analyzes the received visual data using the OpenCV library. This involves detecting the performer's skeletal structure and movements from the video data and obtaining data to evaluate their performance technique. The input is visual data, and the output is the analyzed motion information.

[0827] Step 3: The server analyzes the auditory data using the Librosa library. It extracts musical features, such as tempo and pitch, from the audio data and generates data to evaluate musicality based on these features. The input is auditory data, and the output is the analyzed musical features.

[0828] Step 4: The server uses an emotion analysis engine, such as Hume AI, to analyze visual data and obtain information about the user's emotions. Specifically, it identifies the emotional state during performance from facial expressions and body movements. The input is visual data, and the output is data indicating the emotional state.

[0829] Step 5: The server integrates the previously acquired behavioral information, musical characteristics, and emotional information to calculate an overall musicality score. This score reflects the user's musical skills and emotional expression abilities. The input is the data obtained in Steps 2 through 4, and the output is the overall musicality score.

[0830] Step 6: The server compares the user's desired musicality with the overall musicality score and generates a list to recommend the most suitable musical collaborators. The inputs are the overall musicality score and the user's desired musicality, and the output is a list of collaborators. This list is provided to the user's terminal, enabling matching with appropriate collaborators.

[0831] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0832] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0833] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0834] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0835] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0836] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0837] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0838] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0839] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0840] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0841] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0842] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0843] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0844] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0845] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0846] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0847] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0848] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0849] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0850] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0851] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0852] The following is further disclosed regarding the embodiments described above.

[0853] (Claim 1)

[0854] A means for receiving and saving video data recorded by a user during a performance,

[0855] An analysis means for analyzing movements related to musical instrument performance from the video data,

[0856] An acoustic analysis means for analyzing sound from the moving image data and evaluating its musicality,

[0857] A method for calculating a comprehensive musicality score by integrating the results of video analysis and acoustic analysis,

[0858] A means of comparing the user's desired musical style with the overall musicality score and recommending appropriate music group members,

[0859] A system that includes this.

[0860] (Claim 2)

[0861] The system according to claim 1, which detects the skeleton from the moving image data and analyzes the finger movements of the performer.

[0862] (Claim 3)

[0863] The system according to claim 1, which generates a list of music group members based on the overall musicality score and displays the list on the user's terminal.

[0864] "Example 1"

[0865] (Claim 1)

[0866] A means of receiving and storing multimedia data recorded by a user,

[0867] Processing means for analyzing performance operations from the multimedia data,

[0868] An acoustic processing means for analyzing acoustic information from the multimedia data and evaluating musical characteristics,

[0869] A means for calculating a musicality index by integrating the results of performance motion analysis and sound processing,

[0870] A means of comparing the musical characteristics desired by the user with the musicality index and recommending appropriate musical collaborators,

[0871] A system that includes this.

[0872] (Claim 2)

[0873] The system according to claim 1, which detects body structure from multimedia data and analyzes the finger movements of the performer.

[0874] (Claim 3)

[0875] The system according to claim 1, which generates a list of musical collaborators based on the musicality index and displays the list on the user's terminal.

[0876] "Application Example 1"

[0877] (Claim 1)

[0878] A device that receives and stores video data recorded by a user during a performance,

[0879] A device for analyzing movements related to performance from the video data,

[0880] A device for analyzing sound from the moving image data and evaluating its musicality,

[0881] A device that integrates the results of video analysis and acoustic analysis to calculate musicality evaluation,

[0882] A device that compares the user's desired musical style with an evaluation of that musical style and recommends an appropriate collaborator,

[0883] A device for presenting the collaborator,

[0884] A system that includes this.

[0885] (Claim 2)

[0886] The system according to claim 1, which detects body structure from the moving image data and analyzes the performer's technique.

[0887] (Claim 3)

[0888] The system according to claim 1, which generates a list of collaborators based on the musicality evaluation and displays the list on the user's device.

[0889] "Example 2 of combining an emotion engine"

[0890] (Claim 1)

[0891] A means of receiving and saving video and video materials recorded by the user during a performance,

[0892] An analysis means for analyzing movements related to musical instrument performance from the video data,

[0893] An acoustic analysis means for analyzing sound from the video data and evaluating its musicality,

[0894] A method for calculating an overall musicality evaluation by integrating the results of video analysis and acoustic analysis,

[0895] A means of comparing the user's desired musical style with the overall musicality evaluation and recommending appropriate members for a musical group,

[0896] A means for acquiring emotional data from video and adjusting the overall musicality evaluation based on said emotional data,

[0897] A means for generating a list of members of a musical group based on a adjusted overall musicality evaluation and displaying the list on the user's device,

[0898] A system that includes this.

[0899] (Claim 2)

[0900] The system according to claim 1, which detects the skeleton from the video data and analyzes the finger movements of the performer.

[0901] (Claim 3)

[0902] The system according to claim 1, which adjusts the list of members of a music group based on the emotional data and taking into account the user's emotional state.

[0903] "Application example 2 when combining with an emotional engine"

[0904] (Claim 1)

[0905] A means for receiving and storing visual and auditory data recorded by the user during a performance,

[0906] An analytical means for analyzing movements related to playing a musical instrument from the visual and auditory data,

[0907] An acoustic analysis means for analyzing sound from the visual and auditory data and evaluating its musicality,

[0908] A means for analyzing emotional states from the visual and auditory data and obtaining emotional information,

[0909] A means for calculating a comprehensive musicality score by integrating the results of video analysis, acoustic analysis, and emotional information,

[0910] A means of comparing the user's desired musical style with the overall musicality score and recommending appropriate musical collaborators,

[0911] A system that includes this.

[0912] (Claim 2)

[0913] The system according to claim 1, which detects the structure of the body from the visual and auditory data and analyzes the finger movements of the performer.

[0914] (Claim 3)

[0915] The system according to claim 1, which generates a list of musical collaborators based on the overall musicality score and emotional information, and displays the list on the user's output device. [Explanation of Symbols]

[0916] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving and saving video data recorded by a user during a performance, An analysis means for analyzing movements related to musical instrument performance from the video data, An acoustic analysis means for analyzing sound from the moving image data and evaluating its musicality, A method for calculating a comprehensive musicality score by integrating the results of video analysis and acoustic analysis, A means of comparing the user's desired musical style with the overall musicality score and recommending appropriate music group members, A system that includes this.

2. The system according to claim 1, which detects the skeleton from the moving image data and analyzes the finger movements of the performer.

3. The system according to claim 1, which generates a list of music group members based on the overall musicality score and displays the list on the user's terminal.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A