System

The system addresses the challenge of lacking effective feedback for skill improvement by preprocessing and analyzing audio and image data with a generative AI model, offering users actionable insights and secure, continuous learning opportunities.

JP2026019883APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121631
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Individuals seeking to improve their music or art skills face challenges in receiving appropriate personal feedback and guidance, as traditional methods are often expensive and lack effective means for evaluating and improving their skills.

Method used

A technology improvement support system that processes audio and image data using a generative AI model to provide accurate feedback by performing preprocessing such as noise removal, volume adjustment, background removal, and color correction, and generating specific improvements based on analysis results.

Benefits of technology

Enables efficient and economical skill improvement by providing high-quality, specific feedback that users can easily understand and apply, while ensuring secure data transmission and continuous model improvement through partnerships with educational institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019883000001_ABST
    Figure 2026019883000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A technical improvement support system comprising: means for receiving audio or video; means for pre-processing the audio or video; means for analyzing the pre-processed audio or video using a generative AI model; means for generating specific improvements based on the analyzing; and means for providing the improvements to a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] To improve skills in music, art, etc., one usually needs to attend expensive professional lessons or specialized schools. However, it is difficult to receive appropriate personal feedback to effectively evaluate and improve one's skills, and there are limited means to do so. This invention provides a means to efficiently and economically improve one's skills. [Means for solving the problem]

[0005] To solve this problem, the present invention provides a technology improvement support system including a means for receiving audio data or image data, a means for performing preprocessing, a means for analyzing the data using a generative AI model, a means for generating specific improvements based on the analysis results, and a means for providing the generated improvements to a user. In the case of audio data, noise removal and volume adjustment are performed, and in the case of image data, background removal and color correction are performed, making it possible to provide more accurate feedback.

[0006] "Audio data" refers to sound information, such as a recorded human voice or the sound of a musical instrument, expressed in digital form.

[0007] "Image data" refers to visual information expressed in digital form, such as paintings, photographs, and digital art.

[0008] "Preprocessing" refers to the initial processing carried out to improve the accuracy of the data. In the case of audio data, this includes noise removal and volume adjustment, while in the case of image data, this includes background removal and color correction.

[0009] A "generative AI model" is an algorithm or system that uses machine learning and artificial intelligence techniques to analyze and generate data.

[0010] "Analysis" is the process of extracting and evaluating data characteristics and problems, and providing relevant information.

[0011] "Areas for improvement" refers to specific advice or suggestions for correction or improvement presented based on the analysis results.

[0012] "User" refers to an individual or organization that uses this technology improvement support system to upload audio and image data and receive feedback.

[0013] "Providing" means communicating the generated feedback information to the user and making it available for the user to use. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] The present invention relates to a technology improvement support system using voice data and image data. How the present invention is put into practice will be described below in detail.

[0036] First, the user launches a dedicated application on their smartphone or PC. Through the application, the user creates audio data of their own performance or image data of a painting they have drawn. For example, if the user is singing, they record the audio using the smartphone's microphone. If they are painting, they use the smartphone's camera to take a photo of the work.

[0037] The device then uploads the audio or image data to a server via the application. The HTTPS protocol is used for this upload, ensuring secure transmission of the data to the server. The received data is then tagged with the user's identification information, ensuring accurate feedback to the user.

[0038] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0039] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0040] Once the generative AI model has completed its analysis, the server generates specific improvements based on the results. For example, feedback such as "certain notes are off" or "the rhythm is inconsistent" or "the color scheme is biased" or "the composition is unstable" may be included. The server also provides this feedback in written form, making it easy for users to understand. At the same time, if the data is audio, a corrected version of the audio is also generated.

[0041] Finally, the server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed within the application. The user can learn how to improve their skills based on the specific feedback and use it for their next practice.

[0042] In addition, the server will partner with relevant professional and educational institutions to share data and retrain the AI ​​model to improve its accuracy and increase feedback variations, thereby continuously improving the system's effectiveness and reliability and providing a higher level of support to users.

[0043] The above is an embodiment of the present invention, showing the specific operation of a system for supporting individual skill improvement using audio data and image data.

[0044] The processing flow will be explained below.

[0045] Step 1: Enter your data

[0046] The user starts a dedicated application using a smartphone or a PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image).

[0047] Step 2: Upload your data

[0048] The device uploads the recorded or captured audio and image data to a server, using the HTTPS protocol for secure communication.

[0049] Step 3: Receiving the data

[0050] The server receives the voice data and image data sent from the terminal, as well as the user's identification information.

[0051] Step 4: Preprocessing the audio data

[0052] The server performs noise reduction and volume adjustment on the received audio data, using FFT (Fast Fourier Transform) and noise reduction algorithms.

[0053] Step 5: Preprocessing the image data

[0054] The server performs background removal and color correction on the received image data, using image processing libraries such as OpenCV.

[0055] Step 6: Analysis by generative AI model

[0056] The server inputs preprocessed audio or image data into the generative AI model and analyzes the data. For audio data, it evaluates pitch, rhythm, and vocal accuracy, while for image data, it evaluates composition, color usage, and texture.

[0057] Step 7: Generate improvements

[0058] The server generates specific improvements based on the analysis results obtained from the generative AI model, such as noting that certain notes are off, the rhythm is inconsistent, the color scheme is biased, or the composition is unstable.

[0059] Step 8: Generate corrected audio and video

[0060] The server corrects the audio data and image data as necessary and generates corrected data. In the case of audio data, it creates an audio file with corrected pitch and rhythm.

[0061] Step 9: Provide feedback

[0062] The server sends the generated feedback and correction data back to the user, using notifications to allow the user to view the feedback within the application.

[0063] Step 10: Review feedback

[0064] Users can review the feedback provided within the application, understand specific areas for improvement, and apply the corrective data to their next practice.

[0065] Step 11: Business partnerships and data utilization

[0066] The server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources to retrain the generative AI model and improve the accuracy of the system.

[0067] The above is a description of the specific operations for each processing step. Through this series of processes, users can improve their skills efficiently and effectively.

[0068] Example 1

[0069] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0070] Conventional systems that support individual skill improvement using voice and image data suffer from insufficient preprocessing and analysis, resulting in poor quality feedback provided to users and making it difficult to identify specific areas for improvement. There are also security concerns regarding data transmission. Furthermore, because analysis requires specialized knowledge, users often fail to achieve the results they intended. There is a need for a system that can solve these issues and enable users to truly experience improvements in their skills.

[0071] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0072] In this invention, the server includes: a means for a user to create audio data or image data; a means for securely transmitting the audio data or image data; a means for preprocessing the audio data, such as noise reduction and volume adjustment, or background removal and color correction for the image data; a means for analyzing the preprocessed data using a generative AI model; a means for generating specific improvements based on the analysis results; and a means for providing the improvements to the user. This allows users to receive high-quality, specific feedback, making it easier for them to realize their technical improvements. Furthermore, secure data transmission also improves reliability in terms of security.

[0073] "Means for users to create audio data or image data" refers to a function that allows users to record or take pictures using a dedicated application.

[0074] The "means for securely transmitting the audio data or image data" refers to a function for encrypting data using a security protocol (for example, HTTPS) and securely transmitting the data to a server.

[0075] "Means for performing pre-processing such as noise removal and volume adjustment for audio data, or background removal and color correction for image data" refers to a function that performs processing to improve the quality of data using algorithms and libraries such as FFT for audio data and OpenCV for image data.

[0076] "Means for analyzing preprocessed data using a generative AI model" refers to the function of evaluating preprocessed audio data or image data using an AI model that has previously trained on expert feedback data.

[0077] "Means for generating specific improvements based on the analysis results" refers to a function that generates specific improvements that the user should make (for example, pitch discrepancies or unstable composition) based on the analysis results of the generation AI model.

[0078] "Means for providing the user with the improvements" refers to a function for returning specific improvements obtained as a result of the analysis to the user as text or correction data and notifying them via the application.

[0079] The present invention relates to a technology improvement support system using voice data and image data. How the present invention is put into practice will be described below in detail.

[0080] First, the user launches a dedicated application on their smartphone or PC. Through the application, the user can record audio data of their own performance or create image data of a painting they have drawn. For example, if the user is singing, they can record the audio using the smartphone's microphone. If they are painting, they can take a photo of the work using the smartphone's camera.

[0081] The device then uploads the audio or image data to a server via the application. The HTTPS protocol is used for this upload, ensuring secure transmission of the data to the server. The received data is then tagged with the user's identification information, ensuring accurate feedback to the user.

[0082] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0083] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0084] Once the analysis is complete, the server generates specific improvements based on the results. For example, feedback such as "certain notes are off" or "the rhythm is inconsistent" or "the color scheme is biased" or "the composition is unstable" may be included. The server also provides this feedback in written form, making it easy for users to understand. At the same time, if the data is audio, a corrected version of the audio is also generated.

[0085] Finally, the server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed within the application. The user can learn how to improve their skills based on the specific feedback and use it for their next practice.

[0086] In addition, the server will partner with relevant professional and educational institutions to share data and retrain the AI ​​model to improve its accuracy and increase feedback variations, thereby continuously improving the system's effectiveness and reliability and providing a higher level of support to users.

[0087] Examples of specific examples and prompts

[0088] Example 1: Analyzing audio data

[0089] Users record their own singing on their smartphones and upload the audio data to a server via the application.

[0090] The device securely transmits the audio data to the server using the HTTPS protocol.

[0091] The server uses FFT to remove noise and adjust the volume, while the generative AI model analyzes pitch and rhythm and generates feedback that a particular note is out of tune.

[0092] The user will receive feedback from the application that indicates that a particular note is off, and can use this feedback to improve their next practice.

[0093] Example prompt:

[0094] Analyze my singing performance. I provide audio data. You rate this audio for pitch and rhythmic accuracy and suggest areas for improvement.

[0095] Example 2: Image data analysis

[0096] The user takes a photo of the painting they have created with their smartphone and uploads the image data to the server via the application.

[0097] The device securely transmits the image data to the server using the HTTPS protocol.

[0098] The server uses OpenCV to remove the background of the image and correct the color tone, and the generative AI model analyzes the composition and color usage and generates feedback such as "the color scheme is biased."

[0099] The user can check the feedback in the application that the color scheme is biased and use it to create their next work.

[0100] Example prompt:

[0101] Please analyze a painting. Provide image data. Evaluate the composition and color usage of the work, and suggest areas for improvement.

[0102] The above is an embodiment of the present invention, showing the specific operation of a system for supporting users in improving their individual skills using audio data and image data.

[0103] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0104] Step 1:

[0105] A user launches a dedicated application using a smartphone or PC. The user creates the audio data they want to record or the image data they want to capture. For example, if the user wants to sing, they can use the application's recording function to record the audio. For image data, they can use the application's camera function to take a photo of a painting. The input data can be audio or image.

[0106] Step 2:

[0107] After the user finishes recording or shooting, they check the data in the application and press the save button. This saves the audio or image data to the device. The input data is the audio or image that the user checked and edited, and the saved audio or image file is obtained as output data.

[0108] Step 3:

[0109] The device uploads the saved audio and image data to the server using the HTTPS protocol. The user's identification information is also sent, so accurate feedback can be returned to the user. The input data is the audio file, image file, and user identification information, and is uploaded to the server based on this information.

[0110] Step 4:

[0111] The server preprocesses the received audio and image data. For audio data, it uses FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, it uses OpenCV to remove the background and correct the color tone. The input data is an audio file or an image file, and based on this, data processing such as noise removal, volume adjustment, background removal, and color correction is performed. The output data is a preprocessed audio file or image file.

[0112] Step 5:

[0113] The preprocessed data is input into a generative AI model stored on the server. The generative AI model has previously studied feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, and in the case of image data, it evaluates composition, color usage, and texture. The input data are preprocessed audio and image files, and the output data are the analysis results.

[0114] Step 6:

[0115] The server generates specific improvements based on the analysis results of the generative AI model. For example, it evaluates audio such as "certain pitches are off" or "rhythm is inconsistent," and images such as "color scheme is biased" or "composition is unstable." The input data is the analysis results, and the output data is generated in document format as specific improvements.

[0116] Step 7:

[0117] The server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed by the user within the application. The input data is the generated feedback information, which is used to notify the user. The output data is the specific feedback the user receives.

[0118] Step 8:

[0119] Users can check the provided feedback and use it to improve their skills. They can use the feedback to improve their next practice or work and improve their skills. The input data is feedback information, and the output data is specific actions that will lead to the user's skill improvement.

[0120] (Application example 1)

[0121] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0122] Conventional technology improvement support systems have problems with the insufficient accuracy and speed of analysis and feedback of voice or image data. Another issue is that the devices users can use are limited, making it difficult to provide convenience across a variety of devices. Furthermore, there are insufficient means for providing specific improvements based on the analysis results, and users often lack guidance on how to actually implement improvements.

[0123] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0124] In this invention, the server includes means for receiving voice data or image data, means for preprocessing the voice data or image data, means for analyzing the preprocessed data using a generative AI model, means for generating specific improvements based on the analysis results, means for providing the improvements to the user, means for uploading data to a cloud server and receiving feedback, and means for installation on a smartphone or head-mounted display. This enables highly accurate and rapid analysis and feedback of improvements. Furthermore, the system allows users to use the system on a variety of devices, greatly improving convenience. Furthermore, providing specific improvements allows users to obtain useful guidance for actually improving their performance.

[0125] definition statement

[0126] "Audio data" refers to recorded audio stored in digital format.

[0127] "Image data" refers to data that is a captured or generated image stored in digital format.

[0128] "Preprocessing" refers to initial data processing such as noise removal and color correction to improve the quality of audio and image data.

[0129] A "generative AI model" is an artificial intelligence model that learns using large amounts of data and analyzes and evaluates audio and images.

[0130] A "cloud server" is a remote server that stores, analyzes, and provides data over the Internet.

[0131] A "smartphone" is a portable information terminal with advanced computing power and multiple functions.

[0132] A "head-mounted display" is a device that is worn on the head and displays a display directly within the field of view.

[0133] "User" means a person or entity that uses the system to upload audio or image data and receive feedback.

[0134] "Preprocessed data" refers to audio data or image data that has undergone preprocessing such as noise removal and color correction.

[0135] "Specific improvements" are clear suggestions and advice for improving the user's performance based on the analysis of the generative AI model.

[0136] MODE FOR CARRYING OUT THE INVENTION

[0137] This invention relates to a technology improvement support system using voice and image data. This system receives voice or image data, preprocesses them, analyzes them using a generative AI model, and provides specific feedback on improvements to the user. The following describes how this system is implemented.

[0138] First, the user launches a dedicated application using a smartphone or head-mounted display. The user then creates audio data by recording their own performance or image data by photographing it. For example, if the user is singing, they can record the audio using the smartphone's microphone, or if they are painting, they can take a photo of the artwork using the smartphone's camera.

[0139] The device then uploads the audio or image data to a cloud server via the application. The HTTPS protocol is used for uploading, and the data is securely sent to the server. The server then adds user identification information to the received data, ensuring appropriate feedback.

[0140] The server preprocesses the received audio and image data. For audio data, the server uses FFT (Fast Fourier Transform) and filtering algorithms to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and correct color. This preprocessing ensures that subsequent analysis is performed with high accuracy.

[0141] The preprocessed data is input into a generative AI model stored on the server. The generative AI model is trained based on feedback data collected from music and art experts, enabling highly accurate analysis. Audio data is evaluated for pitch, rhythm, and vocal accuracy, while image data is evaluated for composition, color usage, and texture.

[0142] Once the analysis is complete, the server generates specific improvements based on the results. These improvements include feedback such as "certain notes are off" or "the rhythm is inconsistent," as well as "the color scheme is biased" or "the composition is unstable." Corrected audio and images may also be generated.

[0143] Finally, the server sends the generated feedback information back to the user. The feedback is notified to the user through the application and can be viewed within the application. The user can learn strategies to improve their skills based on this feedback and use it for their next practice.

[0144] As a specific example, analysis is performed by inputting the following prompt sentence into the generative AI model.

[0145] Example of prompt for audio data:

[0146] Please analyze the following audio data and indicate areas for improvement in pitch and rhythm:

[0147] [Audio data]

[0148] Example prompt for image data:

[0149] Analyze the image below and suggest improvements to the color scheme and composition:

[0150] [Image data]

[0151] This system enables accurate and rapid analysis and feedback on improvements, allows users to use the system across a variety of devices, and provides specific improvements, providing useful guidance for users to actually improve their performance.

[0152] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0153] System processing steps

[0154] Step 1:

[0155] The user launches a dedicated application using a smartphone or head-mounted display. The user then creates audio data by recording their own performance or image data by taking photographs. The input is raw audio or image data, obtained by recording or taking photographs. The output is audio or image data stored within the device.

[0156] Step 2:

[0157] The device uploads audio or image data to the cloud server through the application. The HTTPS protocol is used during uploading to ensure data security. The input is the audio or image data stored in the device, and the output is the data uploaded to the cloud server.

[0158] Step 3:

[0159] The server preprocesses the received data. For audio data, FFT (Fast Fourier Transform) and filtering algorithms are used to remove noise and adjust the volume. For image data, OpenCV is used to remove background and correct color. The input is raw data uploaded to the cloud server, and the output is preprocessed data.

[0160] Step 4:

[0161] The server inputs the preprocessed data into the generative AI model. The generative AI model is trained based on feedback data collected in advance from experts, enabling highly accurate analysis. The input is preprocessed data, and the output is the analysis results. In the case of audio data, pitch, rhythm, and vocal accuracy are evaluated, while in the case of image data, composition, color usage, and texture are evaluated.

[0162] Step 5:

[0163] The server generates specific improvements based on the analysis results. This includes feedback such as "certain pitches are off" or "the rhythm is inconsistent," as well as "the color scheme is biased" or "the composition is unstable." The input is the analysis results obtained from the generative AI model, and the output is specific feedback on improvements.

[0164] Step 6:

[0165] The server sends the generated feedback information back to the user. The feedback is notified through the application and can be viewed by the user within the application. The input is specific feedback for improvement, and the output is notification and display to the user.

[0166] This system allows users to learn strategies for improving their own performance based on specific areas for improvement and apply them to their next practice.In addition, the system enables highly accurate and rapid analysis, and can be used on a variety of devices, greatly improving convenience.

[0167] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0168] This invention combines an emotion engine with a technology improvement support system that uses voice data and image data. How this invention is put into practice will be described below in detail.

[0169] First, a user launches a dedicated application on a smartphone or PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image). For example, if a user sings, they use the smartphone's microphone to record the audio. If it is a painting, they use the smartphone's camera to take a photo of the work.

[0170] The device then uploads the audio or image data to a server, using the HTTPS protocol for secure communication. The received data is tagged with the user's identification information, ensuring accurate feedback to the user.

[0171] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0172] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0173] Once the generative AI model has completed its analysis, the server uses an emotion engine to recognize the user's emotions. The server analyzes the voice tone and patterns from the voice data to identify the user's emotional state (e.g., joy, sadness, anger, etc.). The server also analyzes facial expressions and composition from the image data to identify the user's emotional state.

[0174] The server generates specific improvements based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, in addition to pointing out issues such as "certain pitches are off," "rhythms are inconsistent," "color schemes are biased," and "composition is unstable," the server also provides advice based on the user's emotional state. For example, if the emotion recognition result is "nervous," specific advice such as "how to speak in a relaxed manner" is provided.

[0175] The server then corrects the audio data and image data as necessary and generates corrected data. In the case of audio data, it creates an audio file with corrected pitch and rhythm.

[0176] Finally, the server sends the generated feedback and correction data back to the user. The feedback and correction data are notified via the application and can be viewed within the application. The user can understand specific improvements and emotion-based advice and use it for their next practice.

[0177] Additionally, the server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources to retrain the generative AI model and emotion engine, thereby improving the accuracy of the system and providing a higher level of support to users.

[0178] The above describes the mode for carrying out the present invention, and shows the specific operation of a system for supporting individual skill improvement using voice data and image data. By combining it with an emotion engine, feedback is provided that takes into account the user's emotional state, and more effective skill improvement can be expected.

[0179] The processing flow will be explained below.

[0180] Step 1: Enter your data

[0181] A user uses a smartphone or PC to launch a dedicated application and record or photograph audio data (e.g., singing voice) or image data (e.g., painting image). For example, when recording singing voice, the user uses the smartphone's microphone and presses the recording start button within the application. In the case of a painting, the user takes a photo of the work using the smartphone's camera.

[0182] Step 2: Upload your data

[0183] The device uploads recorded or captured audio and image data to a server via the application. Secure communication is performed using the HTTPS protocol. When uploading data, user identification information is also sent.

[0184] Step 3: Receiving the data

[0185] The server receives the voice data or image data sent from the terminal, and then classifies and stores the data appropriately based on the user identification information.

[0186] Step 4: Preprocessing the audio data

[0187] The server performs noise reduction and volume adjustment on the received audio data. Specifically, it applies a noise reduction algorithm using FFT (Fast Fourier Transform) to remove background noise. If the volume is not consistent, it also applies a volume adjustment algorithm.

[0188] Step 5: Preprocessing the image data

[0189] The server performs background removal and color correction on the received image data. For background removal, it uses an image processing library such as OpenCV, and for color correction, it uses an automatic white balance adjustment algorithm.

[0190] Step 6: Analysis by generative AI model

[0191] The server inputs preprocessed audio or image data into the generative AI model, which analyzes the data. For audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy. For image data, the AI ​​model evaluates composition, color, and texture.

[0192] Step 7: Emotion Recognition with the Emotion Engine

[0193] After obtaining the analysis results from the generative AI model, the server uses an emotion engine to recognize the user's emotions. For voice data, it analyzes the voice tone and patterns to identify the user's emotional state (e.g., joy, sadness, anger, etc.). For image data, it uses facial expression analysis algorithms and composition analysis to identify the emotional state.

[0194] Step 8: Generate recommendations based on improvements and sentiment

[0195] The server generates specific improvements based on the analysis results and emotion recognition results obtained from the generative AI model. For example, it points out issues such as "certain pitches are off," "rhythms are inconsistent," "color schemes are biased," and "composition is unstable," and provides advice based on the emotional state (e.g., "how to speak in a relaxed manner").

[0196] Step 9: Generate corrected audio and images

[0197] The server generates modified audio and images based on the original audio or image data. In the case of audio data, it creates an audio file with corrected pitch and rhythm. In the case of image data, it creates an image file with corrected composition and color tone.

[0198] Step 10: Provide feedback

[0199] The server returns the generated feedback information and correction data to the user, who is notified of the feedback and correction data using a notification function so that the user can check it within the application.

[0200] Step 11: Review feedback

[0201] Users can review the feedback and correction data generated within the application, learn specific improvements and emotion-based advice to improve their next practice.

[0202] Step 12: Business partnerships and data utilization

[0203] The server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources, and retrain the generative AI model and emotion engine to improve the accuracy of the system, thereby providing a higher level of support to users.

[0204] This concludes the explanation of the specific operations for each processing step. This series of processes allows users to efficiently and effectively improve their skills, and by adding emotion recognition, they can receive more personalized feedback.

[0205] Example 2

[0206] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0207] Conventional skill improvement support systems do not take into account the user's emotional state when analyzing voice and image data, resulting in feedback that is not adapted to the user's psychological state. Furthermore, the feedback provided is often vague, making it unclear how users should use it to improve their skills. Furthermore, data correction is often done manually, resulting in low efficiency.

[0208] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0209] In this invention, the server includes means for receiving voice data or image data, means for preprocessing the voice data or image data, means for analyzing the preprocessed data using a generative AI model, means for generating specific improvements based on the analysis results and emotion recognition results from an emotion engine, means for providing the user with advice corresponding to the improvements and the user's emotional state, means for correcting the data based on the provided improvements and advice, and means for returning the corrected data and feedback information to the user. This not only supports the user's technical improvement, but also enables highly accurate feedback that takes into account the user's emotional state at the time. Furthermore, the automatic data correction function enables efficient technical improvement support.

[0210] "Audio data" means sound signals recorded using a microphone or other sound recording device.

[0211] "Image data" means signals containing visual information captured using a camera or other image capturing device.

[0212] "Preprocessing" refers to processing carried out to improve the quality of data before data analysis, and includes noise removal and color correction.

[0213] A "generative AI model" refers to an artificial intelligence algorithm that is trained in advance on a large dataset and analyzes and evaluates audio and image data.

[0214] "Emotion Engine" means a software or hardware component for recognizing a user's emotional state from audio and / or visual data.

[0215] "Specific improvements" refers to specific changes or corrections that users should make to improve the technology, based on the analysis results of the generative AI model and the recognition results of the emotion engine.

[0216] "Advice" means guidance or advice to a User based on the results of the generative AI model and emotion engine.

[0217] "Data Correction" means any changes or improvements to the data based on analysis and feedback.

[0218] "Feedback Information" means information, including analytical results and improvements derived from the results of the Generative AI Model and Emotion Engine.

[0219] This invention combines an emotion engine with a technology improvement support system that uses voice data and image data. How this invention is put into practice will now be described in detail.

[0220] A user uses a smartphone or PC to launch a dedicated application and record or photograph audio data (e.g., singing voice) or image data (e.g., painting image). For example, if a user is singing, they can record "Happy Birthday" using the smartphone's microphone. If a user is taking a picture of a painting, they can use the smartphone's camera to photograph a landscape.

[0221] The device uploads the collected audio or image data to a server. This is done securely using the HTTPS protocol. The received data is tagged with the user's identification information, ensuring accurate feedback is sent back to the user.

[0222] The server preprocesses the received audio or image data. For audio data, it uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For example, it filters out background noise and equalizes the overall volume. For image data, it uses image processing libraries such as OpenCV to remove the background and perform color correction. For example, it converts the background of an image to black and white and extracts only the main parts.

[0223] The preprocessed data is input into a generative AI model stored on the server. This generative AI model has previously learned from feedback data from music and art experts, allowing it to perform highly accurate analysis. In the case of audio data, the generative AI model evaluates the accuracy of pitch, rhythm, and vocalization. For example, it may evaluate the data as "the C4 note is low" or "the rhythm is unstable." In the case of image data, it evaluates the composition, color usage, and texture. For example, it may evaluate the data as "the color scheme is unbalanced" or "the composition is biased."

[0224] Once the analysis by the generative AI model is complete, the server uses an emotion engine to analyze the user's emotions. The tone and patterns of the voice data are analyzed to identify the user's emotional state (e.g., joy, tension, relief, etc.). The facial expressions and composition of the image data are also analyzed to similarly identify the emotional state. For example, if the voice tone is bright and tense, it is judged to be "joy," while if the voice is low and the tone is unstable, it is judged to be "tension."

[0225] The server generates specific areas for improvement and advice based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, specific suggestions such as "Your C4 note is low, so you need to practice producing an A4 note accurately" and emotion-based advice such as "If you're nervous, practice vocal exercises that incorporate deep breathing to relax" are provided.

[0226] Furthermore, the server corrects the audio and image data as needed. In the case of audio data, it generates an audio file with corrected pitch and rhythm. For example, it corrects the pitch and rhythm of the singing data of "Happy Birthday" to bring it closer to the correct pitch and rhythm. In the case of image data, it readjusts the color tone and composition.

[0227] Finally, the server sends the generated feedback information and correction data back to the user, who can then review the provided feedback and correction data through a dedicated application, allowing the user to understand specific improvements and emotion-based advice and use it to improve their skills.

[0228] An example of a prompt sentence for voice data (singing voice):

[0229] "Please analyze my recording of "Happy Birthday" and let me know how I can improve it."

[0230] For image data (pictorial images):

[0231] "Please give me feedback on the composition and color scheme of this landscape painting."

[0232] In this way, users can receive feedback and advice on specific techniques to improve their skills and use it in their practice and creative work.

[0233] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0234] Step 1:

[0235] The user starts a dedicated application on their smartphone or PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image). Specifically, the user uses the smartphone's microphone to sing "Happy Birthday" and record the audio. This generates input audio data or image data, which then becomes the input for the next step.

[0236] Step 2:

[0237] The device uploads the collected voice or image data to a server. Secure communication is performed using the HTTPS protocol. The data is accompanied by the user's identification information, ensuring accurate feedback to the user. The input is voice or image data, and the output is the data transferred to the server.

[0238] Step 3:

[0239] The server preprocesses the received audio or image data. For audio data, it uses FFT (Fast Fourier Transform) to remove noise and adjust the volume. Specifically, it filters out background noise and keeps the volume constant. For image data, it uses an image processing library such as OpenCV to remove the background and correct the color tone. The preprocessed data is the input data, and the results of this processing are input to the next step.

[0240] Step 4:

[0241] The server inputs preprocessed audio or image data into the generative AI model. The generative AI model has previously studied expert feedback data, allowing it to perform highly accurate analysis. In the case of audio data, the generative AI model evaluates pitch, rhythm, and accuracy of pronunciation. Specifically, it evaluates whether "the C4 note is low" or "the rhythm is unstable." In the case of image data, it evaluates composition, color usage, and texture. This becomes the analyzed data output.

[0242] Step 5:

[0243] After the server has completed data analysis using the generative AI model, it uses an emotion engine to analyze the user's emotions. In the case of voice data, it analyzes the voice tone and patterns to identify the user's emotional state. For example, it may determine that "if the voice tone is bright and high-tension, it is 'joy'" or "if the voice tone is low and unstable, it is 'tension'." In the case of image data, it analyzes facial expressions and composition to similarly recognize the emotional state. This results in the user's emotion recognition results being output.

[0244] Step 6:

[0245] The server generates specific areas for improvement and advice based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, technical advice such as "Your C4 note is low, so you need to practice producing an A4 accurately" and emotional advice such as "If you're nervous, practice vocalization by incorporating deep breathing to relax" are generated. These are output as areas for improvement and advice.

[0246] Step 7:

[0247] The server corrects the audio data and image data as necessary. In the case of audio data, it generates an audio file with corrected pitch and rhythm. For example, it corrects the pitch and rhythm of the singing data of "Happy Birthday" to bring it closer to the correct pitch and rhythm. In the case of image data, it readjusts the color tone and composition. This is output as corrected data.

[0248] Step 8:

[0249] The server sends the generated feedback information and correction data back to the user, who can then review the provided feedback and correction data through a dedicated application. This allows the user to understand specific improvements and emotional advice, which can be used to improve their skills. This is the final output.

[0250] (Application example 2)

[0251] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0252] In conventional skill improvement support systems, it was difficult to simultaneously provide feedback that took into account the user's emotional state along with specific points for improvement in the user's skills. Furthermore, while real-time technical guidance is required, particularly in actual workplaces such as factories, there was a lack of a means to provide immediate feedback and support the improvement of workers' skills.

[0253] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving audio data or image data, means for preprocessing the audio data or the image data, and means for analyzing the preprocessed data using a generative AI model. This includes means for capturing work video and audio in real time using the camera and microphone of the smart glasses and uploading the data to the server, means for generating feedback in the server based on work improvement points and emotions and displaying it on the display of the smart glasses, means for generating specific improvement points based on the analysis results, and means for providing the improvement points to the user. This makes it possible to provide specific feedback in real time based on the user's technical improvement points and emotions.

[0254] "Audio Data" means information for recording, transmitting, and analyzing audio in digital or analog form.

[0255] "Image data" refers to information for recording, transmitting, and analyzing visual information in digital or analog format.

[0256] "Preprocessing" refers to the initial processing steps taken to convert data into a form suitable for analysis.

[0257] A "generative AI model" is an artificial intelligence system that uses pre-trained data to evaluate and generate new data.

[0258] The "emotion engine" is a system for analyzing and identifying a user's emotional state from voice and image data.

[0259] "Real-time capture" means recording and processing current events and actions as they occur.

[0260] "Smart glasses" are wearable devices that incorporate sensors such as a camera and microphone and provide visual information to the user.

[0261] "Feedback" refers to suggestions and advice for improvement provided based on analysis results and evaluations.

[0262] "Server" means a computer system established for the purpose of processing, storing, and managing data.

[0263] This invention is a system that provides real-time technical guidance to factory workers while they are working using smart glasses. The roles of the server, terminal, and user are as follows:

[0264] Server Roles

[0265] The server uses the following hardware and software to process and analyze data:

[0266] Hardware: A server with a powerful processor and plenty of memory

[0267] software:

[0268] OpenCV: Image processing library

[0269] SciPy: an audio processing library

[0270] Transformers: A library of generative AI models for sentiment analysis

[0271] HTTPS protocol: secure data communication

[0272] The data processing flow is as follows:

[0273] 1. Receive audio or image data transmitted from the smart glasses.

[0274] 2. Preprocess the received data to remove noise, adjust the volume, remove background, and correct color.

[0275] 3. The pre-processed data is fed into a generative AI model for technical analysis.

[0276] 4. Analyze the emotional state of the worker using an emotion engine.

[0277] 5. Generate specific improvements based on these results.

[0278] 6. Generate feedback based on improvements and emotions and display it on the smart glasses display.

[0279] Device Role

[0280] The device (smart glasses) captures and transmits data using the following hardware and software:

[0281] Hardware:

[0282] Camera: Capture footage of your work

[0283] Microphone: Captures audio

[0284] Display: Show feedback

[0285] software:

[0286] Real-time capture and data transmission applications

[0287] The process on the terminal is as follows:

[0288] 1. Use a camera and microphone to capture work video and audio in real time.

[0289] 2. Upload the captured data to the server using the HTTPS protocol.

[0290] User Roles

[0291] The user (worker) wears the smart glasses and performs the work. The user's operation procedure is as follows.

[0292] 1. Put on the smart glasses and start working.

[0293] 2. The camera and microphone in the smart glasses capture the work situation and send it to the server.

[0294] 3. The feedback sent back from the server is viewed on the smart glasses display.

[0295] 4. Correct and improve your work based on feedback.

[0296] Specific examples

[0297] Example: When a factory worker assembles parts, the camera in the smart glasses captures the work process and the microphone records the work instructions. The data is sent to a server, and feedback such as "Part A is facing the wrong way" or "Please work more relaxedly" is displayed in real time based on technical evaluation and emotional state analysis.

[0298] Prompt Sentence Examples

[0299] Image data: Video of parts assembly

[0300] Audio data: "Place part A on part B."

[0301] Issue: Offer advice on how to improve their work and their emotions.

[0302] In this way, users can receive feedback based on their work evaluation and emotions in real time, allowing them to improve their skills efficiently.

[0303] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0304] Step 1:

[0305] The user puts on the smart glasses and starts working. The smart glasses capture the work video with a camera and the audio with a microphone. These data (input) are used to record the work progress and audio instructions in real time. The real-time captured video data and audio data are generated as output.

[0306] Step 2:

[0307] The device (smart glasses) securely uploads the captured video and audio data to the server using the HTTPS protocol. This process ensures the data is securely transmitted to the server. The input is the captured video and audio data, and the output is the video and audio data transmitted to the server.

[0308] Step 3:

[0309] The server preprocesses the received video and audio data. Specifically, it performs noise reduction and volume adjustment on the audio data, and background removal and color correction on the video data. The input is the video and audio data sent to the server, and the output is the preprocessed data. This is done using libraries such as OpenCV and SciPy.

[0310] Step 4:

[0311] The server inputs the preprocessed data into the generative AI model and performs technical analysis. The generative AI model analyzes the input data based on the data it has previously learned. The input is the preprocessed data, and the output is the analysis results (specific technical improvements).

[0312] Step 5:

[0313] The server uses an emotion engine to analyze the user's emotional state from voice tone and visual expressions. The input is preprocessed data, and the output is the recognition result of the emotional state. This is done based on voice tone, patterns, visual expressions, etc.

[0314] Step 6:

[0315] The server generates specific improvement points and feedback according to emotions based on the results of technical analysis and emotion recognition. The inputs are the analysis results and emotion recognition results, and the output is feedback to be provided to the user (advice on how to improve work and respond to emotions).

[0316] Step 7:

[0317] The server sends the generated feedback to the device (smart glasses) using the HTTPS protocol, so that the user can check the feedback in real time. The input is the feedback data, and the output is the feedback sent to the smart glasses.

[0318] Step 8:

[0319] The terminal (smart glasses) displays the received feedback on its display. The user checks the feedback in real time and corrects and improves their work. The input is the feedback sent from the server, and the output is the feedback displayed on the smart glasses display.

[0320] In this way, data is captured, transferred, processed, and feedback is provided at each step, allowing users to receive real-time technical improvement and emotionally sensitive guidance.

[0321] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0322] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0323] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0324] [Second embodiment]

[0325] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0326] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0327] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0328] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0329] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0330] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0331] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0332] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0333] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0334] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0335] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0336] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0337] The present invention relates to a technology improvement support system using voice data and image data. How the present invention is put into practice will be described below in detail.

[0338] First, the user launches a dedicated application on their smartphone or PC. Through the application, the user creates audio data of their own performance or image data of a painting they have drawn. For example, if the user is singing, they record the audio using the smartphone's microphone. If they are painting, they use the smartphone's camera to take a photo of the work.

[0339] The device then uploads the audio or image data to a server via the application. The HTTPS protocol is used for this upload, ensuring secure transmission of the data to the server. The received data is then tagged with the user's identification information, ensuring accurate feedback to the user.

[0340] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0341] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0342] Once the generative AI model has completed its analysis, the server generates specific improvements based on the results. For example, feedback such as "certain notes are off" or "the rhythm is inconsistent" or "the color scheme is biased" or "the composition is unstable" may be included. The server also provides this feedback in written form, making it easy for users to understand. At the same time, if the data is audio, a corrected version of the audio is also generated.

[0343] Finally, the server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed within the application. The user can learn how to improve their skills based on the specific feedback and use it for their next practice.

[0344] In addition, the server will partner with relevant professional and educational institutions to share data and retrain the AI ​​model to improve its accuracy and increase feedback variations, thereby continuously improving the system's effectiveness and reliability and providing a higher level of support to users.

[0345] The above is an embodiment of the present invention, showing the specific operation of a system for supporting individual skill improvement using audio data and image data.

[0346] The processing flow will be explained below.

[0347] Step 1: Enter your data

[0348] The user starts a dedicated application using a smartphone or a PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image).

[0349] Step 2: Upload your data

[0350] The device uploads the recorded or captured audio and image data to a server, using the HTTPS protocol for secure communication.

[0351] Step 3: Receiving the data

[0352] The server receives the voice data and image data sent from the terminal, as well as the user's identification information.

[0353] Step 4: Preprocessing the audio data

[0354] The server performs noise reduction and volume adjustment on the received audio data, using FFT (Fast Fourier Transform) and noise reduction algorithms.

[0355] Step 5: Preprocessing the image data

[0356] The server performs background removal and color correction on the received image data, using image processing libraries such as OpenCV.

[0357] Step 6: Analysis by generative AI model

[0358] The server inputs preprocessed audio or image data into the generative AI model and analyzes the data. For audio data, it evaluates pitch, rhythm, and vocal accuracy, while for image data, it evaluates composition, color usage, and texture.

[0359] Step 7: Generate improvements

[0360] The server generates specific improvements based on the analysis results obtained from the generative AI model, such as noting that certain notes are off, the rhythm is inconsistent, the color scheme is biased, or the composition is unstable.

[0361] Step 8: Generate corrected audio and video

[0362] The server corrects the audio data and image data as necessary and generates corrected data. In the case of audio data, it creates an audio file with corrected pitch and rhythm.

[0363] Step 9: Provide feedback

[0364] The server sends the generated feedback and correction data back to the user, using notifications to allow the user to view the feedback within the application.

[0365] Step 10: Review feedback

[0366] Users can review the feedback provided within the application, understand specific areas for improvement, and apply the corrective data to their next practice.

[0367] Step 11: Business partnerships and data utilization

[0368] The server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources to retrain the generative AI model and improve the accuracy of the system.

[0369] The above is a description of the specific operations for each processing step. Through this series of processes, users can improve their skills efficiently and effectively.

[0370] Example 1

[0371] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0372] Conventional systems that support individual skill improvement using voice and image data suffer from insufficient preprocessing and analysis, resulting in poor quality feedback provided to users and making it difficult to identify specific areas for improvement. There are also security concerns regarding data transmission. Furthermore, because analysis requires specialized knowledge, users often fail to achieve the results they intended. There is a need for a system that can solve these issues and enable users to truly experience improvements in their skills.

[0373] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0374] In this invention, the server includes: a means for a user to create audio data or image data; a means for securely transmitting the audio data or image data; a means for preprocessing the audio data, such as noise reduction and volume adjustment, or background removal and color correction for the image data; a means for analyzing the preprocessed data using a generative AI model; a means for generating specific improvements based on the analysis results; and a means for providing the improvements to the user. This allows users to receive high-quality, specific feedback, making it easier for them to realize their technical improvements. Furthermore, secure data transmission also improves reliability in terms of security.

[0375] "Means for users to create audio data or image data" refers to a function that allows users to record or take pictures using a dedicated application.

[0376] The "means for securely transmitting the audio data or image data" refers to a function for encrypting data using a security protocol (for example, HTTPS) and securely transmitting the data to a server.

[0377] "Means for performing pre-processing such as noise removal and volume adjustment for audio data, or background removal and color correction for image data" refers to a function that performs processing to improve the quality of data using algorithms and libraries such as FFT for audio data and OpenCV for image data.

[0378] "Means for analyzing preprocessed data using a generative AI model" refers to the function of evaluating preprocessed audio data or image data using an AI model that has previously trained on expert feedback data.

[0379] "Means for generating specific improvements based on the analysis results" refers to a function that generates specific improvements that the user should make (for example, pitch discrepancies or unstable composition) based on the analysis results of the generation AI model.

[0380] "Means for providing the user with the improvements" refers to a function for returning specific improvements obtained as a result of the analysis to the user as text or correction data and notifying them via the application.

[0381] The present invention relates to a technology improvement support system using voice data and image data. How the present invention is put into practice will be described below in detail.

[0382] First, the user launches a dedicated application on their smartphone or PC. Through the application, the user can record audio data of their own performance or create image data of a painting they have drawn. For example, if the user is singing, they can record the audio using the smartphone's microphone. If they are painting, they can take a photo of the work using the smartphone's camera.

[0383] The device then uploads the audio or image data to a server via the application. The HTTPS protocol is used for this upload, ensuring secure transmission of the data to the server. The received data is then tagged with the user's identification information, ensuring accurate feedback to the user.

[0384] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0385] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0386] Once the analysis is complete, the server generates specific improvements based on the results. For example, feedback such as "certain notes are off" or "the rhythm is inconsistent" or "the color scheme is biased" or "the composition is unstable" may be included. The server also provides this feedback in written form, making it easy for users to understand. At the same time, if the data is audio, a corrected version of the audio is also generated.

[0387] Finally, the server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed within the application. The user can learn how to improve their skills based on the specific feedback and use it for their next practice.

[0388] In addition, the server will partner with relevant professional and educational institutions to share data and retrain the AI ​​model to improve its accuracy and increase feedback variations, thereby continuously improving the system's effectiveness and reliability and providing a higher level of support to users.

[0389] Examples of specific examples and prompts

[0390] Example 1: Analyzing audio data

[0391] Users record their own singing on their smartphones and upload the audio data to a server via the application.

[0392] The device securely transmits the audio data to the server using the HTTPS protocol.

[0393] The server uses FFT to remove noise and adjust the volume, while the generative AI model analyzes pitch and rhythm and generates feedback that a particular note is out of tune.

[0394] The user will receive feedback from the application that indicates that a particular note is off, and can use this feedback to improve their next practice.

[0395] Example prompt:

[0396] Analyze my singing performance. I provide audio data. You rate this audio for pitch and rhythmic accuracy and suggest areas for improvement.

[0397] Example 2: Image data analysis

[0398] The user takes a photo of the painting they have created with their smartphone and uploads the image data to the server via the application.

[0399] The device securely transmits the image data to the server using the HTTPS protocol.

[0400] The server uses OpenCV to remove the background of the image and correct the color tone, and the generative AI model analyzes the composition and color usage and generates feedback such as "the color scheme is biased."

[0401] The user can check the feedback in the application that the color scheme is biased and use it to create their next work.

[0402] Example prompt:

[0403] Please analyze a painting. Provide image data. Evaluate the composition and color usage of the work, and suggest areas for improvement.

[0404] The above is an embodiment of the present invention, showing the specific operation of a system for supporting users in improving their individual skills using audio data and image data.

[0405] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0406] Step 1:

[0407] A user launches a dedicated application using a smartphone or PC. The user creates the audio data they want to record or the image data they want to capture. For example, if the user wants to sing, they can use the application's recording function to record the audio. For image data, they can use the application's camera function to take a photo of a painting. The input data can be audio or image.

[0408] Step 2:

[0409] After the user finishes recording or shooting, they check the data in the application and press the save button. This saves the audio or image data to the device. The input data is the audio or image that the user checked and edited, and the saved audio or image file is obtained as output data.

[0410] Step 3:

[0411] The device uploads the saved audio and image data to the server using the HTTPS protocol. The user's identification information is also sent, so accurate feedback can be returned to the user. The input data is the audio file, image file, and user identification information, and is uploaded to the server based on this information.

[0412] Step 4:

[0413] The server preprocesses the received audio and image data. For audio data, it uses FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, it uses OpenCV to remove the background and correct the color tone. The input data is an audio file or an image file, and based on this, data processing such as noise removal, volume adjustment, background removal, and color correction is performed. The output data is a preprocessed audio file or image file.

[0414] Step 5:

[0415] The preprocessed data is input into a generative AI model stored on the server. The generative AI model has previously studied feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, and in the case of image data, it evaluates composition, color usage, and texture. The input data are preprocessed audio and image files, and the output data are the analysis results.

[0416] Step 6:

[0417] The server generates specific improvements based on the analysis results of the generative AI model. For example, it evaluates audio such as "certain pitches are off" or "rhythm is inconsistent," and images such as "color scheme is biased" or "composition is unstable." The input data is the analysis results, and the output data is generated in document format as specific improvements.

[0418] Step 7:

[0419] The server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed by the user within the application. The input data is the generated feedback information, which is used to notify the user. The output data is the specific feedback the user receives.

[0420] Step 8:

[0421] Users can check the provided feedback and use it to improve their skills. They can use the feedback to improve their next practice or work and improve their skills. The input data is feedback information, and the output data is specific actions that will lead to the user's skill improvement.

[0422] (Application example 1)

[0423] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0424] Conventional technology improvement support systems have problems with the insufficient accuracy and speed of analysis and feedback of voice or image data. Another issue is that the devices users can use are limited, making it difficult to provide convenience across a variety of devices. Furthermore, there are insufficient means for providing specific improvements based on the analysis results, and users often lack guidance on how to actually implement improvements.

[0425] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0426] In this invention, the server includes means for receiving voice data or image data, means for preprocessing the voice data or image data, means for analyzing the preprocessed data using a generative AI model, means for generating specific improvements based on the analysis results, means for providing the improvements to the user, means for uploading data to a cloud server and receiving feedback, and means for installation on a smartphone or head-mounted display. This enables highly accurate and rapid analysis and feedback of improvements. Furthermore, the system allows users to use the system on a variety of devices, greatly improving convenience. Furthermore, providing specific improvements allows users to obtain useful guidance for actually improving their performance.

[0427] definition statement

[0428] "Audio data" refers to recorded audio stored in digital format.

[0429] "Image data" refers to data that is a captured or generated image stored in digital format.

[0430] "Preprocessing" refers to initial data processing such as noise removal and color correction to improve the quality of audio and image data.

[0431] A "generative AI model" is an artificial intelligence model that learns using large amounts of data and analyzes and evaluates audio and images.

[0432] A "cloud server" is a remote server that stores, analyzes, and provides data over the Internet.

[0433] A "smartphone" is a portable information terminal with advanced computing power and multiple functions.

[0434] A "head-mounted display" is a device that is worn on the head and displays a display directly within the field of view.

[0435] "User" means a person or entity that uses the system to upload audio or image data and receive feedback.

[0436] "Preprocessed data" refers to audio data or image data that has undergone preprocessing such as noise removal and color correction.

[0437] "Specific improvements" are clear suggestions and advice for improving the user's performance based on the analysis of the generative AI model.

[0438] MODE FOR CARRYING OUT THE INVENTION

[0439] This invention relates to a technology improvement support system using voice and image data. This system receives voice or image data, preprocesses them, analyzes them using a generative AI model, and provides specific feedback on improvements to the user. The following describes how this system is implemented.

[0440] First, the user launches a dedicated application using a smartphone or head-mounted display. The user then creates audio data by recording their own performance or image data by photographing it. For example, if the user is singing, they can record the audio using the smartphone's microphone, or if they are painting, they can take a photo of the artwork using the smartphone's camera.

[0441] The device then uploads the audio or image data to a cloud server via the application. The HTTPS protocol is used for uploading, and the data is securely sent to the server. The server then adds user identification information to the received data, ensuring appropriate feedback.

[0442] The server preprocesses the received audio and image data. For audio data, the server uses FFT (Fast Fourier Transform) and filtering algorithms to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and correct color. This preprocessing ensures that subsequent analysis is performed with high accuracy.

[0443] The preprocessed data is input into a generative AI model stored on the server. The generative AI model is trained based on feedback data collected from music and art experts, enabling highly accurate analysis. Audio data is evaluated for pitch, rhythm, and vocal accuracy, while image data is evaluated for composition, color usage, and texture.

[0444] Once the analysis is complete, the server generates specific improvements based on the results. These improvements include feedback such as "certain notes are off" or "the rhythm is inconsistent," as well as "the color scheme is biased" or "the composition is unstable." Corrected audio and images may also be generated.

[0445] Finally, the server sends the generated feedback information back to the user. The feedback is notified to the user through the application and can be viewed within the application. The user can learn strategies to improve their skills based on this feedback and use it for their next practice.

[0446] As a specific example, analysis is performed by inputting the following prompt sentence into the generative AI model.

[0447] Example of prompt for audio data:

[0448] Please analyze the following audio data and indicate areas for improvement in pitch and rhythm:

[0449] [Audio data]

[0450] Example prompt for image data:

[0451] Analyze the image below and suggest improvements to the color scheme and composition:

[0452] [Image data]

[0453] This system enables accurate and rapid analysis and feedback on improvements, allows users to use the system across a variety of devices, and provides specific improvements, providing useful guidance for users to actually improve their performance.

[0454] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0455] System processing steps

[0456] Step 1:

[0457] The user launches a dedicated application using a smartphone or head-mounted display. The user then creates audio data by recording their own performance or image data by taking photographs. The input is raw audio or image data, obtained by recording or taking photographs. The output is audio or image data stored within the device.

[0458] Step 2:

[0459] The device uploads audio or image data to the cloud server through the application. The HTTPS protocol is used during uploading to ensure data security. The input is the audio or image data stored in the device, and the output is the data uploaded to the cloud server.

[0460] Step 3:

[0461] The server preprocesses the received data. For audio data, FFT (Fast Fourier Transform) and filtering algorithms are used to remove noise and adjust the volume. For image data, OpenCV is used to remove background and correct color. The input is raw data uploaded to the cloud server, and the output is preprocessed data.

[0462] Step 4:

[0463] The server inputs the preprocessed data into the generative AI model. The generative AI model is trained based on feedback data collected in advance from experts, enabling highly accurate analysis. The input is preprocessed data, and the output is the analysis results. In the case of audio data, pitch, rhythm, and vocal accuracy are evaluated, while in the case of image data, composition, color usage, and texture are evaluated.

[0464] Step 5:

[0465] The server generates specific improvements based on the analysis results. This includes feedback such as "certain pitches are off" or "the rhythm is inconsistent," as well as "the color scheme is biased" or "the composition is unstable." The input is the analysis results obtained from the generative AI model, and the output is specific feedback on improvements.

[0466] Step 6:

[0467] The server sends the generated feedback information back to the user. The feedback is notified through the application and can be viewed by the user within the application. The input is specific feedback for improvement, and the output is notification and display to the user.

[0468] This system allows users to learn strategies for improving their own performance based on specific areas for improvement and apply them to their next practice.In addition, the system enables highly accurate and rapid analysis, and can be used on a variety of devices, greatly improving convenience.

[0469] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0470] This invention combines an emotion engine with a technology improvement support system that uses voice data and image data. How this invention is put into practice will be described below in detail.

[0471] First, a user launches a dedicated application on a smartphone or PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image). For example, if a user sings, they use the smartphone's microphone to record the audio. If it is a painting, they use the smartphone's camera to take a photo of the work.

[0472] The device then uploads the audio or image data to a server, using the HTTPS protocol for secure communication. The received data is tagged with the user's identification information, ensuring accurate feedback to the user.

[0473] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0474] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0475] Once the generative AI model has completed its analysis, the server uses an emotion engine to recognize the user's emotions. The server analyzes the voice tone and patterns from the voice data to identify the user's emotional state (e.g., joy, sadness, anger, etc.). The server also analyzes facial expressions and composition from the image data to identify the user's emotional state.

[0476] The server generates specific improvements based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, in addition to pointing out issues such as "certain pitches are off," "rhythms are inconsistent," "color schemes are biased," and "composition is unstable," the server also provides advice based on the user's emotional state. For example, if the emotion recognition result is "nervous," specific advice such as "how to speak in a relaxed manner" is provided.

[0477] The server then corrects the audio data and image data as necessary and generates corrected data. In the case of audio data, it creates an audio file with corrected pitch and rhythm.

[0478] Finally, the server sends the generated feedback and correction data back to the user. The feedback and correction data are notified via the application and can be viewed within the application. The user can understand specific improvements and emotion-based advice and use it for their next practice.

[0479] Additionally, the server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources to retrain the generative AI model and emotion engine, thereby improving the accuracy of the system and providing a higher level of support to users.

[0480] The above describes the mode for carrying out the present invention, and shows the specific operation of a system for supporting individual skill improvement using voice data and image data. By combining it with an emotion engine, feedback is provided that takes into account the user's emotional state, and more effective skill improvement can be expected.

[0481] The processing flow will be explained below.

[0482] Step 1: Enter your data

[0483] A user uses a smartphone or PC to launch a dedicated application and record or photograph audio data (e.g., singing voice) or image data (e.g., painting image). For example, when recording singing voice, the user uses the smartphone's microphone and presses the recording start button within the application. In the case of a painting, the user takes a photo of the work using the smartphone's camera.

[0484] Step 2: Upload your data

[0485] The device uploads recorded or captured audio and image data to a server via the application. Secure communication is performed using the HTTPS protocol. When uploading data, user identification information is also sent.

[0486] Step 3: Receiving the data

[0487] The server receives the voice data or image data sent from the terminal, and then classifies and stores the data appropriately based on the user identification information.

[0488] Step 4: Preprocessing the audio data

[0489] The server performs noise reduction and volume adjustment on the received audio data. Specifically, it applies a noise reduction algorithm using FFT (Fast Fourier Transform) to remove background noise. If the volume is not consistent, it also applies a volume adjustment algorithm.

[0490] Step 5: Preprocessing the image data

[0491] The server performs background removal and color correction on the received image data. For background removal, it uses an image processing library such as OpenCV, and for color correction, it uses an automatic white balance adjustment algorithm.

[0492] Step 6: Analysis by generative AI model

[0493] The server inputs preprocessed audio or image data into the generative AI model, which analyzes the data. For audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy. For image data, the AI ​​model evaluates composition, color, and texture.

[0494] Step 7: Emotion Recognition with the Emotion Engine

[0495] After obtaining the analysis results from the generative AI model, the server uses an emotion engine to recognize the user's emotions. For voice data, it analyzes the voice tone and patterns to identify the user's emotional state (e.g., joy, sadness, anger, etc.). For image data, it uses facial expression analysis algorithms and composition analysis to identify the emotional state.

[0496] Step 8: Generate recommendations based on improvements and sentiment

[0497] The server generates specific improvements based on the analysis results and emotion recognition results obtained from the generative AI model. For example, it points out issues such as "certain pitches are off," "rhythms are inconsistent," "color schemes are biased," and "composition is unstable," and provides advice based on the emotional state (e.g., "how to speak in a relaxed manner").

[0498] Step 9: Generate corrected audio and images

[0499] The server generates modified audio and images based on the original audio or image data. In the case of audio data, it creates an audio file with corrected pitch and rhythm. In the case of image data, it creates an image file with corrected composition and color tone.

[0500] Step 10: Provide feedback

[0501] The server returns the generated feedback information and correction data to the user, who is notified of the feedback and correction data using a notification function so that the user can check it within the application.

[0502] Step 11: Review feedback

[0503] Users can review the feedback and correction data generated within the application, learn specific improvements and emotion-based advice to improve their next practice.

[0504] Step 12: Business partnerships and data utilization

[0505] The server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources, and retrain the generative AI model and emotion engine to improve the accuracy of the system, thereby providing a higher level of support to users.

[0506] This concludes the explanation of the specific operations for each processing step. This series of processes allows users to efficiently and effectively improve their skills, and by adding emotion recognition, they can receive more personalized feedback.

[0507] Example 2

[0508] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0509] Conventional skill improvement support systems do not take into account the user's emotional state when analyzing voice and image data, resulting in feedback that is not adapted to the user's psychological state. Furthermore, the feedback provided is often vague, making it unclear how users should use it to improve their skills. Furthermore, data correction is often done manually, resulting in low efficiency.

[0510] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0511] In this invention, the server includes means for receiving voice data or image data, means for preprocessing the voice data or image data, means for analyzing the preprocessed data using a generative AI model, means for generating specific improvements based on the analysis results and emotion recognition results from an emotion engine, means for providing the user with advice corresponding to the improvements and the user's emotional state, means for correcting the data based on the provided improvements and advice, and means for returning the corrected data and feedback information to the user. This not only supports the user's technical improvement, but also enables highly accurate feedback that takes into account the user's emotional state at the time. Furthermore, the automatic data correction function enables efficient technical improvement support.

[0512] "Audio data" means sound signals recorded using a microphone or other sound recording device.

[0513] "Image data" means signals containing visual information captured using a camera or other image capturing device.

[0514] "Preprocessing" refers to processing carried out to improve the quality of data before data analysis, and includes noise removal and color correction.

[0515] A "generative AI model" refers to an artificial intelligence algorithm that is trained in advance on a large dataset and analyzes and evaluates audio and image data.

[0516] "Emotion Engine" means a software or hardware component for recognizing a user's emotional state from audio and / or visual data.

[0517] "Specific improvements" refers to specific changes or corrections that users should make to improve the technology, based on the analysis results of the generative AI model and the recognition results of the emotion engine.

[0518] "Advice" means guidance or advice to a User based on the results of the generative AI model and emotion engine.

[0519] "Data Correction" means any changes or improvements to the data based on analysis and feedback.

[0520] "Feedback Information" means information, including analytical results and improvements derived from the results of the Generative AI Model and Emotion Engine.

[0521] This invention combines an emotion engine with a technology improvement support system that uses voice data and image data. How this invention is put into practice will now be described in detail.

[0522] A user uses a smartphone or PC to launch a dedicated application and record or photograph audio data (e.g., singing voice) or image data (e.g., painting image). For example, if a user is singing, they can record "Happy Birthday" using the smartphone's microphone. If a user is taking a picture of a painting, they can use the smartphone's camera to photograph a landscape.

[0523] The device uploads the collected audio or image data to a server. This is done securely using the HTTPS protocol. The received data is tagged with the user's identification information, ensuring accurate feedback is sent back to the user.

[0524] The server preprocesses the received audio or image data. For audio data, it uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For example, it filters out background noise and equalizes the overall volume. For image data, it uses image processing libraries such as OpenCV to remove the background and perform color correction. For example, it converts the background of an image to black and white and extracts only the main parts.

[0525] The preprocessed data is input into a generative AI model stored on the server. This generative AI model has previously learned from feedback data from music and art experts, allowing it to perform highly accurate analysis. In the case of audio data, the generative AI model evaluates the accuracy of pitch, rhythm, and vocalization. For example, it may evaluate the data as "the C4 note is low" or "the rhythm is unstable." In the case of image data, it evaluates the composition, color usage, and texture. For example, it may evaluate the data as "the color scheme is unbalanced" or "the composition is biased."

[0526] Once the analysis by the generative AI model is complete, the server uses an emotion engine to analyze the user's emotions. The tone and patterns of the voice data are analyzed to identify the user's emotional state (e.g., joy, tension, relief, etc.). The facial expressions and composition of the image data are also analyzed to similarly identify the emotional state. For example, if the voice tone is bright and tense, it is judged to be "joy," while if the voice is low and the tone is unstable, it is judged to be "tension."

[0527] The server generates specific areas for improvement and advice based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, specific suggestions such as "Your C4 note is low, so you need to practice producing an A4 note accurately" and emotion-based advice such as "If you're nervous, practice vocal exercises that incorporate deep breathing to relax" are provided.

[0528] Furthermore, the server corrects the audio and image data as needed. In the case of audio data, it generates an audio file with corrected pitch and rhythm. For example, it corrects the pitch and rhythm of the singing data of "Happy Birthday" to bring it closer to the correct pitch and rhythm. In the case of image data, it readjusts the color tone and composition.

[0529] Finally, the server sends the generated feedback information and correction data back to the user, who can then review the provided feedback and correction data through a dedicated application, allowing the user to understand specific improvements and emotion-based advice and use it to improve their skills.

[0530] An example of a prompt sentence for voice data (singing voice):

[0531] "Please analyze my recording of "Happy Birthday" and let me know how I can improve it."

[0532] For image data (pictorial images):

[0533] "Please give me feedback on the composition and color scheme of this landscape painting."

[0534] In this way, users can receive feedback and advice on specific techniques to improve their skills and use it in their practice and creative work.

[0535] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0536] Step 1:

[0537] The user starts a dedicated application on their smartphone or PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image). Specifically, the user uses the smartphone's microphone to sing "Happy Birthday" and record the audio. This generates input audio data or image data, which then becomes the input for the next step.

[0538] Step 2:

[0539] The device uploads the collected voice or image data to a server. Secure communication is performed using the HTTPS protocol. The data is accompanied by the user's identification information, ensuring accurate feedback to the user. The input is voice or image data, and the output is the data transferred to the server.

[0540] Step 3:

[0541] The server preprocesses the received audio or image data. For audio data, it uses FFT (Fast Fourier Transform) to remove noise and adjust the volume. Specifically, it filters out background noise and keeps the volume constant. For image data, it uses an image processing library such as OpenCV to remove the background and correct the color tone. The preprocessed data is the input data, and the results of this processing are input to the next step.

[0542] Step 4:

[0543] The server inputs preprocessed audio or image data into the generative AI model. The generative AI model has previously studied expert feedback data, allowing it to perform highly accurate analysis. In the case of audio data, the generative AI model evaluates pitch, rhythm, and accuracy of pronunciation. Specifically, it evaluates whether "the C4 note is low" or "the rhythm is unstable." In the case of image data, it evaluates composition, color usage, and texture. This becomes the analyzed data output.

[0544] Step 5:

[0545] After the server has completed data analysis using the generative AI model, it uses an emotion engine to analyze the user's emotions. In the case of voice data, it analyzes the voice tone and patterns to identify the user's emotional state. For example, it may determine that "if the voice tone is bright and high-tension, it is 'joy'" or "if the voice tone is low and unstable, it is 'tension'." In the case of image data, it analyzes facial expressions and composition to similarly recognize the emotional state. This results in the user's emotion recognition results being output.

[0546] Step 6:

[0547] The server generates specific areas for improvement and advice based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, technical advice such as "Your C4 note is low, so you need to practice producing an A4 accurately" and emotional advice such as "If you're nervous, practice vocalization by incorporating deep breathing to relax" are generated. These are output as areas for improvement and advice.

[0548] Step 7:

[0549] The server corrects the audio data and image data as necessary. In the case of audio data, it generates an audio file with corrected pitch and rhythm. For example, it corrects the pitch and rhythm of the singing data of "Happy Birthday" to bring it closer to the correct pitch and rhythm. In the case of image data, it readjusts the color tone and composition. This is output as corrected data.

[0550] Step 8:

[0551] The server sends the generated feedback information and correction data back to the user, who can then review the provided feedback and correction data through a dedicated application. This allows the user to understand specific improvements and emotional advice, which can be used to improve their skills. This is the final output.

[0552] (Application example 2)

[0553] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0554] In conventional skill improvement support systems, it was difficult to simultaneously provide feedback that took into account the user's emotional state along with specific points for improvement in the user's skills. Furthermore, while real-time technical guidance is required, particularly in actual workplaces such as factories, there was a lack of a means to provide immediate feedback and support the improvement of workers' skills.

[0555] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving audio data or image data, means for preprocessing the audio data or the image data, and means for analyzing the preprocessed data using a generative AI model. This includes means for capturing work video and audio in real time using the camera and microphone of the smart glasses and uploading the data to the server, means for generating feedback in the server based on work improvement points and emotions and displaying it on the display of the smart glasses, means for generating specific improvement points based on the analysis results, and means for providing the improvement points to the user. This makes it possible to provide specific feedback in real time based on the user's technical improvement points and emotions.

[0556] "Audio Data" means information for recording, transmitting, and analyzing audio in digital or analog form.

[0557] "Image data" refers to information for recording, transmitting, and analyzing visual information in digital or analog format.

[0558] "Preprocessing" refers to the initial processing steps taken to convert data into a form suitable for analysis.

[0559] A "generative AI model" is an artificial intelligence system that uses pre-trained data to evaluate and generate new data.

[0560] The "emotion engine" is a system for analyzing and identifying a user's emotional state from voice and image data.

[0561] "Real-time capture" means recording and processing current events and actions as they occur.

[0562] "Smart glasses" are wearable devices that incorporate sensors such as a camera and microphone and provide visual information to the user.

[0563] "Feedback" refers to suggestions and advice for improvement provided based on analysis results and evaluations.

[0564] "Server" means a computer system established for the purpose of processing, storing, and managing data.

[0565] This invention is a system that provides real-time technical guidance to factory workers while they are working using smart glasses. The roles of the server, terminal, and user are as follows:

[0566] Server Roles

[0567] The server uses the following hardware and software to process and analyze data:

[0568] Hardware: A server with a powerful processor and plenty of memory

[0569] software:

[0570] OpenCV: Image processing library

[0571] SciPy: an audio processing library

[0572] Transformers: A library of generative AI models for sentiment analysis

[0573] HTTPS protocol: secure data communication

[0574] The data processing flow is as follows:

[0575] 1. Receive audio or image data transmitted from the smart glasses.

[0576] 2. Preprocess the received data to remove noise, adjust the volume, remove background, and correct color.

[0577] 3. The pre-processed data is fed into a generative AI model for technical analysis.

[0578] 4. Analyze the emotional state of the worker using an emotion engine.

[0579] 5. Generate specific improvements based on these results.

[0580] 6. Generate feedback based on improvements and emotions and display it on the smart glasses display.

[0581] Device Role

[0582] The device (smart glasses) captures and transmits data using the following hardware and software:

[0583] Hardware:

[0584] Camera: Capture footage of your work

[0585] Microphone: Captures audio

[0586] Display: Show feedback

[0587] software:

[0588] Real-time capture and data transmission applications

[0589] The process on the terminal is as follows:

[0590] 1. Use a camera and microphone to capture work video and audio in real time.

[0591] 2. Upload the captured data to the server using the HTTPS protocol.

[0592] User Roles

[0593] The user (worker) wears the smart glasses and performs the work. The user's operation procedure is as follows.

[0594] 1. Put on the smart glasses and start working.

[0595] 2. The camera and microphone in the smart glasses capture the work situation and send it to the server.

[0596] 3. The feedback sent back from the server is viewed on the smart glasses display.

[0597] 4. Correct and improve your work based on feedback.

[0598] Specific examples

[0599] Example: When a factory worker assembles parts, the camera in the smart glasses captures the work process and the microphone records the work instructions. The data is sent to a server, and feedback such as "Part A is facing the wrong way" or "Please work more relaxedly" is displayed in real time based on technical evaluation and emotional state analysis.

[0600] Prompt Sentence Examples

[0601] Image data: Video of parts assembly

[0602] Audio data: "Place part A on part B."

[0603] Issue: Offer advice on how to improve their work and their emotions.

[0604] In this way, users can receive feedback based on their work evaluation and emotions in real time, allowing them to improve their skills efficiently.

[0605] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0606] Step 1:

[0607] The user puts on the smart glasses and starts working. The smart glasses capture the work video with a camera and the audio with a microphone. These data (input) are used to record the work progress and audio instructions in real time. The real-time captured video data and audio data are generated as output.

[0608] Step 2:

[0609] The device (smart glasses) securely uploads the captured video and audio data to the server using the HTTPS protocol. This process ensures the data is securely transmitted to the server. The input is the captured video and audio data, and the output is the video and audio data transmitted to the server.

[0610] Step 3:

[0611] The server preprocesses the received video and audio data. Specifically, it performs noise reduction and volume adjustment on the audio data, and background removal and color correction on the video data. The input is the video and audio data sent to the server, and the output is the preprocessed data. This is done using libraries such as OpenCV and SciPy.

[0612] Step 4:

[0613] The server inputs the preprocessed data into the generative AI model and performs technical analysis. The generative AI model analyzes the input data based on the data it has previously learned. The input is the preprocessed data, and the output is the analysis results (specific technical improvements).

[0614] Step 5:

[0615] The server uses an emotion engine to analyze the user's emotional state from voice tone and visual expressions. The input is preprocessed data, and the output is the recognition result of the emotional state. This is done based on voice tone, patterns, visual expressions, etc.

[0616] Step 6:

[0617] The server generates specific improvement points and feedback according to emotions based on the results of technical analysis and emotion recognition. The inputs are the analysis results and emotion recognition results, and the output is feedback to be provided to the user (advice on how to improve work and respond to emotions).

[0618] Step 7:

[0619] The server sends the generated feedback to the device (smart glasses) using the HTTPS protocol, so that the user can check the feedback in real time. The input is the feedback data, and the output is the feedback sent to the smart glasses.

[0620] Step 8:

[0621] The terminal (smart glasses) displays the received feedback on its display. The user checks the feedback in real time and corrects and improves their work. The input is the feedback sent from the server, and the output is the feedback displayed on the smart glasses display.

[0622] In this way, data is captured, transferred, processed, and feedback is provided at each step, allowing users to receive real-time technical improvement and emotionally sensitive guidance.

[0623] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0624] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0625] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0626] [Third embodiment]

[0627] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0628] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0629] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0630] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0631] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0632] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0633] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0634] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0635] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0636] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0637] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0638] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0639] The present invention relates to a technology improvement support system using voice data and image data. How the present invention is put into practice will be described below in detail.

[0640] First, the user launches a dedicated application on their smartphone or PC. Through the application, the user creates audio data of their own performance or image data of a painting they have drawn. For example, if the user is singing, they record the audio using the smartphone's microphone. If they are painting, they use the smartphone's camera to take a photo of the work.

[0641] The device then uploads the audio or image data to a server via the application. The HTTPS protocol is used for this upload, ensuring secure transmission of the data to the server. The received data is then tagged with the user's identification information, ensuring accurate feedback to the user.

[0642] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0643] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0644] Once the generative AI model has completed its analysis, the server generates specific improvements based on the results. For example, feedback such as "certain notes are off" or "the rhythm is inconsistent" or "the color scheme is biased" or "the composition is unstable" may be included. The server also provides this feedback in written form, making it easy for users to understand. At the same time, if the data is audio, a corrected version of the audio is also generated.

[0645] Finally, the server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed within the application. The user can learn how to improve their skills based on the specific feedback and use it for their next practice.

[0646] In addition, the server will partner with relevant professional and educational institutions to share data and retrain the AI ​​model to improve its accuracy and increase feedback variations, thereby continuously improving the system's effectiveness and reliability and providing a higher level of support to users.

[0647] The above is an embodiment of the present invention, showing the specific operation of a system for supporting individual skill improvement using audio data and image data.

[0648] The processing flow will be explained below.

[0649] Step 1: Enter your data

[0650] The user starts a dedicated application using a smartphone or a PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image).

[0651] Step 2: Upload your data

[0652] The device uploads the recorded or captured audio and image data to a server, using the HTTPS protocol for secure communication.

[0653] Step 3: Receiving the data

[0654] The server receives the voice data and image data sent from the terminal, as well as the user's identification information.

[0655] Step 4: Preprocessing the audio data

[0656] The server performs noise reduction and volume adjustment on the received audio data, using FFT (Fast Fourier Transform) and noise reduction algorithms.

[0657] Step 5: Preprocessing the image data

[0658] The server performs background removal and color correction on the received image data, using image processing libraries such as OpenCV.

[0659] Step 6: Analysis by generative AI model

[0660] The server inputs preprocessed audio or image data into the generative AI model and analyzes the data. For audio data, it evaluates pitch, rhythm, and vocal accuracy, while for image data, it evaluates composition, color usage, and texture.

[0661] Step 7: Generate improvements

[0662] The server generates specific improvements based on the analysis results obtained from the generative AI model, such as noting that certain notes are off, the rhythm is inconsistent, the color scheme is biased, or the composition is unstable.

[0663] Step 8: Generate corrected audio and video

[0664] The server corrects the audio data and image data as necessary and generates corrected data. In the case of audio data, it creates an audio file with corrected pitch and rhythm.

[0665] Step 9: Provide feedback

[0666] The server sends the generated feedback and correction data back to the user, using notifications to allow the user to view the feedback within the application.

[0667] Step 10: Review feedback

[0668] Users can review the feedback provided within the application, understand specific areas for improvement, and apply the corrective data to their next practice.

[0669] Step 11: Business partnerships and data utilization

[0670] The server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources to retrain the generative AI model and improve the accuracy of the system.

[0671] The above is a description of the specific operations for each processing step. Through this series of processes, users can improve their skills efficiently and effectively.

[0672] Example 1

[0673] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0674] Conventional systems that support individual skill improvement using voice and image data suffer from insufficient preprocessing and analysis, resulting in poor quality feedback provided to users and making it difficult to identify specific areas for improvement. There are also security concerns regarding data transmission. Furthermore, because analysis requires specialized knowledge, users often fail to achieve the results they intended. There is a need for a system that can solve these issues and enable users to truly experience improvements in their skills.

[0675] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0676] In this invention, the server includes: a means for a user to create audio data or image data; a means for securely transmitting the audio data or image data; a means for preprocessing the audio data, such as noise reduction and volume adjustment, or background removal and color correction for the image data; a means for analyzing the preprocessed data using a generative AI model; a means for generating specific improvements based on the analysis results; and a means for providing the improvements to the user. This allows users to receive high-quality, specific feedback, making it easier for them to realize their technical improvements. Furthermore, secure data transmission also improves reliability in terms of security.

[0677] "Means for users to create audio data or image data" refers to a function that allows users to record or take pictures using a dedicated application.

[0678] The "means for securely transmitting the audio data or image data" refers to a function for encrypting data using a security protocol (for example, HTTPS) and securely transmitting the data to a server.

[0679] "Means for performing pre-processing such as noise removal and volume adjustment for audio data, or background removal and color correction for image data" refers to a function that performs processing to improve the quality of data using algorithms and libraries such as FFT for audio data and OpenCV for image data.

[0680] "Means for analyzing preprocessed data using a generative AI model" refers to the function of evaluating preprocessed audio data or image data using an AI model that has previously trained on expert feedback data.

[0681] "Means for generating specific improvements based on the analysis results" refers to a function that generates specific improvements that the user should make (for example, pitch discrepancies or unstable composition) based on the analysis results of the generation AI model.

[0682] "Means for providing the user with the improvements" refers to a function for returning specific improvements obtained as a result of the analysis to the user as text or correction data and notifying them via the application.

[0683] The present invention relates to a technology improvement support system using voice data and image data. How the present invention is put into practice will be described below in detail.

[0684] First, the user launches a dedicated application on their smartphone or PC. Through the application, the user can record audio data of their own performance or create image data of a painting they have drawn. For example, if the user is singing, they can record the audio using the smartphone's microphone. If they are painting, they can take a photo of the work using the smartphone's camera.

[0685] The device then uploads the audio or image data to a server via the application. The HTTPS protocol is used for this upload, ensuring secure transmission of the data to the server. The received data is then tagged with the user's identification information, ensuring accurate feedback to the user.

[0686] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0687] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0688] Once the analysis is complete, the server generates specific improvements based on the results. For example, feedback such as "certain notes are off" or "the rhythm is inconsistent" or "the color scheme is biased" or "the composition is unstable" may be included. The server also provides this feedback in written form, making it easy for users to understand. At the same time, if the data is audio, a corrected version of the audio is also generated.

[0689] Finally, the server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed within the application. The user can learn how to improve their skills based on the specific feedback and use it for their next practice.

[0690] In addition, the server will partner with relevant professional and educational institutions to share data and retrain the AI ​​model to improve its accuracy and increase feedback variations, thereby continuously improving the system's effectiveness and reliability and providing a higher level of support to users.

[0691] Examples of specific examples and prompts

[0692] Example 1: Analyzing audio data

[0693] Users record their own singing on their smartphones and upload the audio data to a server via the application.

[0694] The device securely transmits the audio data to the server using the HTTPS protocol.

[0695] The server uses FFT to remove noise and adjust the volume, while the generative AI model analyzes pitch and rhythm and generates feedback that a particular note is out of tune.

[0696] The user will receive feedback from the application that indicates that a particular note is off, and can use this feedback to improve their next practice.

[0697] Example prompt:

[0698] Analyze my singing performance. I provide audio data. You rate this audio for pitch and rhythmic accuracy and suggest areas for improvement.

[0699] Example 2: Image data analysis

[0700] The user takes a photo of the painting they have created with their smartphone and uploads the image data to the server via the application.

[0701] The device securely transmits the image data to the server using the HTTPS protocol.

[0702] The server uses OpenCV to remove the background of the image and correct the color tone, and the generative AI model analyzes the composition and color usage and generates feedback such as "the color scheme is biased."

[0703] The user can check the feedback in the application that the color scheme is biased and use it to create their next work.

[0704] Example prompt:

[0705] Please analyze a painting. Provide image data. Evaluate the composition and color usage of the work, and suggest areas for improvement.

[0706] The above is an embodiment of the present invention, showing the specific operation of a system for supporting users in improving their individual skills using audio data and image data.

[0707] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0708] Step 1:

[0709] A user launches a dedicated application using a smartphone or PC. The user creates the audio data they want to record or the image data they want to capture. For example, if the user wants to sing, they can use the application's recording function to record the audio. For image data, they can use the application's camera function to take a photo of a painting. The input data can be audio or image.

[0710] Step 2:

[0711] After the user finishes recording or shooting, they check the data in the application and press the save button. This saves the audio or image data to the device. The input data is the audio or image that the user checked and edited, and the saved audio or image file is obtained as output data.

[0712] Step 3:

[0713] The device uploads the saved audio and image data to the server using the HTTPS protocol. The user's identification information is also sent, so accurate feedback can be returned to the user. The input data is the audio file, image file, and user identification information, and is uploaded to the server based on this information.

[0714] Step 4:

[0715] The server preprocesses the received audio and image data. For audio data, it uses FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, it uses OpenCV to remove the background and correct the color tone. The input data is an audio file or an image file, and based on this, data processing such as noise removal, volume adjustment, background removal, and color correction is performed. The output data is a preprocessed audio file or image file.

[0716] Step 5:

[0717] The preprocessed data is input into a generative AI model stored on the server. The generative AI model has previously studied feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, and in the case of image data, it evaluates composition, color usage, and texture. The input data are preprocessed audio and image files, and the output data are the analysis results.

[0718] Step 6:

[0719] The server generates specific improvements based on the analysis results of the generative AI model. For example, it evaluates audio such as "certain pitches are off" or "rhythm is inconsistent," and images such as "color scheme is biased" or "composition is unstable." The input data is the analysis results, and the output data is generated in document format as specific improvements.

[0720] Step 7:

[0721] The server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed by the user within the application. The input data is the generated feedback information, which is used to notify the user. The output data is the specific feedback the user receives.

[0722] Step 8:

[0723] Users can check the provided feedback and use it to improve their skills. They can use the feedback to improve their next practice or work and improve their skills. The input data is feedback information, and the output data is specific actions that will lead to the user's skill improvement.

[0724] (Application example 1)

[0725] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0726] Conventional technology improvement support systems have problems with the insufficient accuracy and speed of analysis and feedback of voice or image data. Another issue is that the devices users can use are limited, making it difficult to provide convenience across a variety of devices. Furthermore, there are insufficient means for providing specific improvements based on the analysis results, and users often lack guidance on how to actually implement improvements.

[0727] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0728] In this invention, the server includes means for receiving voice data or image data, means for preprocessing the voice data or image data, means for analyzing the preprocessed data using a generative AI model, means for generating specific improvements based on the analysis results, means for providing the improvements to the user, means for uploading data to a cloud server and receiving feedback, and means for installation on a smartphone or head-mounted display. This enables highly accurate and rapid analysis and feedback of improvements. Furthermore, the system allows users to use the system on a variety of devices, greatly improving convenience. Furthermore, providing specific improvements allows users to obtain useful guidance for actually improving their performance.

[0729] definition statement

[0730] "Audio data" refers to recorded audio stored in digital format.

[0731] "Image data" refers to data that is a captured or generated image stored in digital format.

[0732] "Preprocessing" refers to initial data processing such as noise removal and color correction to improve the quality of audio and image data.

[0733] A "generative AI model" is an artificial intelligence model that learns using large amounts of data and analyzes and evaluates audio and images.

[0734] A "cloud server" is a remote server that stores, analyzes, and provides data over the Internet.

[0735] A "smartphone" is a portable information terminal with advanced computing power and multiple functions.

[0736] A "head-mounted display" is a device that is worn on the head and displays a display directly within the field of view.

[0737] "User" means a person or entity that uses the system to upload audio or image data and receive feedback.

[0738] "Preprocessed data" refers to audio data or image data that has undergone preprocessing such as noise removal and color correction.

[0739] "Specific improvements" are clear suggestions and advice for improving the user's performance based on the analysis of the generative AI model.

[0740] MODE FOR CARRYING OUT THE INVENTION

[0741] This invention relates to a technology improvement support system using voice and image data. This system receives voice or image data, preprocesses them, analyzes them using a generative AI model, and provides specific feedback on improvements to the user. The following describes how this system is implemented.

[0742] First, the user launches a dedicated application using a smartphone or head-mounted display. The user then creates audio data by recording their own performance or image data by photographing it. For example, if the user is singing, they can record the audio using the smartphone's microphone, or if they are painting, they can take a photo of the artwork using the smartphone's camera.

[0743] The device then uploads the audio or image data to a cloud server via the application. The HTTPS protocol is used for uploading, and the data is securely sent to the server. The server then adds user identification information to the received data, ensuring appropriate feedback.

[0744] The server preprocesses the received audio and image data. For audio data, the server uses FFT (Fast Fourier Transform) and filtering algorithms to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and correct color. This preprocessing ensures that subsequent analysis is performed with high accuracy.

[0745] The preprocessed data is input into a generative AI model stored on the server. The generative AI model is trained based on feedback data collected from music and art experts, enabling highly accurate analysis. Audio data is evaluated for pitch, rhythm, and vocal accuracy, while image data is evaluated for composition, color usage, and texture.

[0746] Once the analysis is complete, the server generates specific improvements based on the results. These improvements include feedback such as "certain notes are off" or "the rhythm is inconsistent," as well as "the color scheme is biased" or "the composition is unstable." Corrected audio and images may also be generated.

[0747] Finally, the server sends the generated feedback information back to the user. The feedback is notified to the user through the application and can be viewed within the application. The user can learn strategies to improve their skills based on this feedback and use it for their next practice.

[0748] As a specific example, analysis is performed by inputting the following prompt sentence into the generative AI model.

[0749] Example of prompt for audio data:

[0750] Please analyze the following audio data and indicate areas for improvement in pitch and rhythm:

[0751] [Audio data]

[0752] Example prompt for image data:

[0753] Analyze the image below and suggest improvements to the color scheme and composition:

[0754] [Image data]

[0755] This system enables accurate and rapid analysis and feedback on improvements, allows users to use the system across a variety of devices, and provides specific improvements, providing useful guidance for users to actually improve their performance.

[0756] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0757] System processing steps

[0758] Step 1:

[0759] The user launches a dedicated application using a smartphone or head-mounted display. The user then creates audio data by recording their own performance or image data by taking photographs. The input is raw audio or image data, obtained by recording or taking photographs. The output is audio or image data stored within the device.

[0760] Step 2:

[0761] The device uploads audio or image data to the cloud server through the application. The HTTPS protocol is used during uploading to ensure data security. The input is the audio or image data stored in the device, and the output is the data uploaded to the cloud server.

[0762] Step 3:

[0763] The server preprocesses the received data. For audio data, FFT (Fast Fourier Transform) and filtering algorithms are used to remove noise and adjust the volume. For image data, OpenCV is used to remove background and correct color. The input is raw data uploaded to the cloud server, and the output is preprocessed data.

[0764] Step 4:

[0765] The server inputs the preprocessed data into the generative AI model. The generative AI model is trained based on feedback data collected in advance from experts, enabling highly accurate analysis. The input is preprocessed data, and the output is the analysis results. In the case of audio data, pitch, rhythm, and vocal accuracy are evaluated, while in the case of image data, composition, color usage, and texture are evaluated.

[0766] Step 5:

[0767] The server generates specific improvements based on the analysis results. This includes feedback such as "certain pitches are off" or "the rhythm is inconsistent," as well as "the color scheme is biased" or "the composition is unstable." The input is the analysis results obtained from the generative AI model, and the output is specific feedback on improvements.

[0768] Step 6:

[0769] The server sends the generated feedback information back to the user. The feedback is notified through the application and can be viewed by the user within the application. The input is specific feedback for improvement, and the output is notification and display to the user.

[0770] This system allows users to learn strategies for improving their own performance based on specific areas for improvement and apply them to their next practice.In addition, the system enables highly accurate and rapid analysis, and can be used on a variety of devices, greatly improving convenience.

[0771] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0772] This invention combines an emotion engine with a technology improvement support system that uses voice data and image data. How this invention is put into practice will be described below in detail.

[0773] First, a user launches a dedicated application on a smartphone or PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image). For example, if a user sings, they use the smartphone's microphone to record the audio. If it is a painting, they use the smartphone's camera to take a photo of the work.

[0774] The device then uploads the audio or image data to a server, using the HTTPS protocol for secure communication. The received data is tagged with the user's identification information, ensuring accurate feedback to the user.

[0775] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0776] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0777] Once the generative AI model has completed its analysis, the server uses an emotion engine to recognize the user's emotions. The server analyzes the voice tone and patterns from the voice data to identify the user's emotional state (e.g., joy, sadness, anger, etc.). The server also analyzes facial expressions and composition from the image data to identify the user's emotional state.

[0778] The server generates specific improvements based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, in addition to pointing out issues such as "certain pitches are off," "rhythms are inconsistent," "color schemes are biased," and "composition is unstable," the server also provides advice based on the user's emotional state. For example, if the emotion recognition result is "nervous," specific advice such as "how to speak in a relaxed manner" is provided.

[0779] The server then corrects the audio data and image data as necessary and generates corrected data. In the case of audio data, it creates an audio file with corrected pitch and rhythm.

[0780] Finally, the server sends the generated feedback and correction data back to the user. The feedback and correction data are notified via the application and can be viewed within the application. The user can understand specific improvements and emotion-based advice and use it for their next practice.

[0781] Additionally, the server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources to retrain the generative AI model and emotion engine, thereby improving the accuracy of the system and providing a higher level of support to users.

[0782] The above describes the mode for carrying out the present invention, and shows the specific operation of a system for supporting individual skill improvement using voice data and image data. By combining it with an emotion engine, feedback is provided that takes into account the user's emotional state, and more effective skill improvement can be expected.

[0783] The processing flow will be explained below.

[0784] Step 1: Enter your data

[0785] A user uses a smartphone or PC to launch a dedicated application and record or photograph audio data (e.g., singing voice) or image data (e.g., painting image). For example, when recording singing voice, the user uses the smartphone's microphone and presses the recording start button within the application. In the case of a painting, the user takes a photo of the work using the smartphone's camera.

[0786] Step 2: Upload your data

[0787] The device uploads recorded or captured audio and image data to a server via the application. Secure communication is performed using the HTTPS protocol. When uploading data, user identification information is also sent.

[0788] Step 3: Receiving the data

[0789] The server receives the voice data or image data sent from the terminal, and then classifies and stores the data appropriately based on the user identification information.

[0790] Step 4: Preprocessing the audio data

[0791] The server performs noise reduction and volume adjustment on the received audio data. Specifically, it applies a noise reduction algorithm using FFT (Fast Fourier Transform) to remove background noise. If the volume is not consistent, it also applies a volume adjustment algorithm.

[0792] Step 5: Preprocessing the image data

[0793] The server performs background removal and color correction on the received image data. For background removal, it uses an image processing library such as OpenCV, and for color correction, it uses an automatic white balance adjustment algorithm.

[0794] Step 6: Analysis by generative AI model

[0795] The server inputs preprocessed audio or image data into the generative AI model, which analyzes the data. For audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy. For image data, the AI ​​model evaluates composition, color, and texture.

[0796] Step 7: Emotion Recognition with the Emotion Engine

[0797] After obtaining the analysis results from the generative AI model, the server uses an emotion engine to recognize the user's emotions. For voice data, it analyzes the voice tone and patterns to identify the user's emotional state (e.g., joy, sadness, anger, etc.). For image data, it uses facial expression analysis algorithms and composition analysis to identify the emotional state.

[0798] Step 8: Generate recommendations based on improvements and sentiment

[0799] The server generates specific improvements based on the analysis results and emotion recognition results obtained from the generative AI model. For example, it points out issues such as "certain pitches are off," "rhythms are inconsistent," "color schemes are biased," and "composition is unstable," and provides advice based on the emotional state (e.g., "how to speak in a relaxed manner").

[0800] Step 9: Generate corrected audio and images

[0801] The server generates modified audio and images based on the original audio or image data. In the case of audio data, it creates an audio file with corrected pitch and rhythm. In the case of image data, it creates an image file with corrected composition and color tone.

[0802] Step 10: Provide feedback

[0803] The server returns the generated feedback information and correction data to the user, who is notified of the feedback and correction data using a notification function so that the user can check it within the application.

[0804] Step 11: Review feedback

[0805] Users can review the feedback and correction data generated within the application, learn specific improvements and emotion-based advice to improve their next practice.

[0806] Step 12: Business partnerships and data utilization

[0807] The server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources, and retrain the generative AI model and emotion engine to improve the accuracy of the system, thereby providing a higher level of support to users.

[0808] This concludes the explanation of the specific operations for each processing step. This series of processes allows users to efficiently and effectively improve their skills, and by adding emotion recognition, they can receive more personalized feedback.

[0809] Example 2

[0810] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0811] Conventional skill improvement support systems do not take into account the user's emotional state when analyzing voice and image data, resulting in feedback that is not adapted to the user's psychological state. Furthermore, the feedback provided is often vague, making it unclear how users should use it to improve their skills. Furthermore, data correction is often done manually, resulting in low efficiency.

[0812] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0813] In this invention, the server includes means for receiving voice data or image data, means for preprocessing the voice data or image data, means for analyzing the preprocessed data using a generative AI model, means for generating specific improvements based on the analysis results and emotion recognition results from an emotion engine, means for providing the user with advice corresponding to the improvements and the user's emotional state, means for correcting the data based on the provided improvements and advice, and means for returning the corrected data and feedback information to the user. This not only supports the user's technical improvement, but also enables highly accurate feedback that takes into account the user's emotional state at the time. Furthermore, the automatic data correction function enables efficient technical improvement support.

[0814] "Audio data" means sound signals recorded using a microphone or other sound recording device.

[0815] "Image data" means signals containing visual information captured using a camera or other image capturing device.

[0816] "Preprocessing" refers to processing carried out to improve the quality of data before data analysis, and includes noise removal and color correction.

[0817] A "generative AI model" refers to an artificial intelligence algorithm that is trained in advance on a large dataset and analyzes and evaluates audio and image data.

[0818] "Emotion Engine" means a software or hardware component for recognizing a user's emotional state from audio and / or visual data.

[0819] "Specific improvements" refers to specific changes or corrections that users should make to improve the technology, based on the analysis results of the generative AI model and the recognition results of the emotion engine.

[0820] "Advice" means guidance or advice to a User based on the results of the generative AI model and emotion engine.

[0821] "Data Correction" means any changes or improvements to the data based on analysis and feedback.

[0822] "Feedback Information" means information, including analytical results and improvements derived from the results of the Generative AI Model and Emotion Engine.

[0823] This invention combines an emotion engine with a technology improvement support system that uses voice data and image data. How this invention is put into practice will now be described in detail.

[0824] A user uses a smartphone or PC to launch a dedicated application and record or photograph audio data (e.g., singing voice) or image data (e.g., painting image). For example, if a user is singing, they can record "Happy Birthday" using the smartphone's microphone. If a user is taking a picture of a painting, they can use the smartphone's camera to photograph a landscape.

[0825] The device uploads the collected audio or image data to a server. This is done securely using the HTTPS protocol. The received data is tagged with the user's identification information, ensuring accurate feedback is sent back to the user.

[0826] The server preprocesses the received audio or image data. For audio data, it uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For example, it filters out background noise and equalizes the overall volume. For image data, it uses image processing libraries such as OpenCV to remove the background and perform color correction. For example, it converts the background of an image to black and white and extracts only the main parts.

[0827] The preprocessed data is input into a generative AI model stored on the server. This generative AI model has previously learned from feedback data from music and art experts, allowing it to perform highly accurate analysis. In the case of audio data, the generative AI model evaluates the accuracy of pitch, rhythm, and vocalization. For example, it may evaluate the data as "the C4 note is low" or "the rhythm is unstable." In the case of image data, it evaluates the composition, color usage, and texture. For example, it may evaluate the data as "the color scheme is unbalanced" or "the composition is biased."

[0828] Once the analysis by the generative AI model is complete, the server uses an emotion engine to analyze the user's emotions. The tone and patterns of the voice data are analyzed to identify the user's emotional state (e.g., joy, tension, relief, etc.). The facial expressions and composition of the image data are also analyzed to similarly identify the emotional state. For example, if the voice tone is bright and tense, it is judged to be "joy," while if the voice is low and the tone is unstable, it is judged to be "tension."

[0829] The server generates specific areas for improvement and advice based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, specific suggestions such as "Your C4 note is low, so you need to practice producing an A4 note accurately" and emotion-based advice such as "If you're nervous, practice vocal exercises that incorporate deep breathing to relax" are provided.

[0830] Furthermore, the server corrects the audio and image data as needed. In the case of audio data, it generates an audio file with corrected pitch and rhythm. For example, it corrects the pitch and rhythm of the singing data of "Happy Birthday" to bring it closer to the correct pitch and rhythm. In the case of image data, it readjusts the color tone and composition.

[0831] Finally, the server sends the generated feedback information and correction data back to the user, who can then review the provided feedback and correction data through a dedicated application, allowing the user to understand specific improvements and emotion-based advice and use it to improve their skills.

[0832] An example of a prompt sentence for voice data (singing voice):

[0833] "Please analyze my recording of "Happy Birthday" and let me know how I can improve it."

[0834] For image data (pictorial images):

[0835] "Please give me feedback on the composition and color scheme of this landscape painting."

[0836] In this way, users can receive feedback and advice on specific techniques to improve their skills and use it in their practice and creative work.

[0837] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0838] Step 1:

[0839] The user starts a dedicated application on their smartphone or PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image). Specifically, the user uses the smartphone's microphone to sing "Happy Birthday" and record the audio. This generates input audio data or image data, which then becomes the input for the next step.

[0840] Step 2:

[0841] The device uploads the collected voice or image data to a server. Secure communication is performed using the HTTPS protocol. The data is accompanied by the user's identification information, ensuring accurate feedback to the user. The input is voice or image data, and the output is the data transferred to the server.

[0842] Step 3:

[0843] The server preprocesses the received audio or image data. For audio data, it uses FFT (Fast Fourier Transform) to remove noise and adjust the volume. Specifically, it filters out background noise and keeps the volume constant. For image data, it uses an image processing library such as OpenCV to remove the background and correct the color tone. The preprocessed data is the input data, and the results of this processing are input to the next step.

[0844] Step 4:

[0845] The server inputs preprocessed audio or image data into the generative AI model. The generative AI model has previously studied expert feedback data, allowing it to perform highly accurate analysis. In the case of audio data, the generative AI model evaluates pitch, rhythm, and accuracy of pronunciation. Specifically, it evaluates whether "the C4 note is low" or "the rhythm is unstable." In the case of image data, it evaluates composition, color usage, and texture. This becomes the analyzed data output.

[0846] Step 5:

[0847] After the server has completed data analysis using the generative AI model, it uses an emotion engine to analyze the user's emotions. In the case of voice data, it analyzes the voice tone and patterns to identify the user's emotional state. For example, it may determine that "if the voice tone is bright and high-tension, it is 'joy'" or "if the voice tone is low and unstable, it is 'tension'." In the case of image data, it analyzes facial expressions and composition to similarly recognize the emotional state. This results in the user's emotion recognition results being output.

[0848] Step 6:

[0849] The server generates specific areas for improvement and advice based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, technical advice such as "Your C4 note is low, so you need to practice producing an A4 accurately" and emotional advice such as "If you're nervous, practice vocalization by incorporating deep breathing to relax" are generated. These are output as areas for improvement and advice.

[0850] Step 7:

[0851] The server corrects the audio data and image data as necessary. In the case of audio data, it generates an audio file with corrected pitch and rhythm. For example, it corrects the pitch and rhythm of the singing data of "Happy Birthday" to bring it closer to the correct pitch and rhythm. In the case of image data, it readjusts the color tone and composition. This is output as corrected data.

[0852] Step 8:

[0853] The server sends the generated feedback information and correction data back to the user, who can then review the provided feedback and correction data through a dedicated application. This allows the user to understand specific improvements and emotional advice, which can be used to improve their skills. This is the final output.

[0854] (Application example 2)

[0855] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0856] In conventional skill improvement support systems, it was difficult to simultaneously provide feedback that took into account the user's emotional state along with specific points for improvement in the user's skills. Furthermore, while real-time technical guidance is required, particularly in actual workplaces such as factories, there was a lack of a means to provide immediate feedback and support the improvement of workers' skills.

[0857] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving audio data or image data, means for preprocessing the audio data or the image data, and means for analyzing the preprocessed data using a generative AI model. This includes means for capturing work video and audio in real time using the camera and microphone of the smart glasses and uploading the data to the server, means for generating feedback in the server based on work improvement points and emotions and displaying it on the display of the smart glasses, means for generating specific improvement points based on the analysis results, and means for providing the improvement points to the user. This makes it possible to provide specific feedback in real time based on the user's technical improvement points and emotions.

[0858] "Audio Data" means information for recording, transmitting, and analyzing audio in digital or analog form.

[0859] "Image data" refers to information for recording, transmitting, and analyzing visual information in digital or analog format.

[0860] "Preprocessing" refers to the initial processing steps taken to convert data into a form suitable for analysis.

[0861] A "generative AI model" is an artificial intelligence system that uses pre-trained data to evaluate and generate new data.

[0862] The "emotion engine" is a system for analyzing and identifying a user's emotional state from voice and image data.

[0863] "Real-time capture" means recording and processing current events and actions as they occur.

[0864] "Smart glasses" are wearable devices that incorporate sensors such as a camera and microphone and provide visual information to the user.

[0865] "Feedback" refers to suggestions and advice for improvement provided based on analysis results and evaluations.

[0866] "Server" means a computer system established for the purpose of processing, storing, and managing data.

[0867] This invention is a system that provides real-time technical guidance to factory workers while they are working using smart glasses. The roles of the server, terminal, and user are as follows:

[0868] Server Roles

[0869] The server uses the following hardware and software to process and analyze data:

[0870] Hardware: A server with a powerful processor and plenty of memory

[0871] software:

[0872] OpenCV: Image processing library

[0873] SciPy: an audio processing library

[0874] Transformers: A library of generative AI models for sentiment analysis

[0875] HTTPS protocol: secure data communication

[0876] The data processing flow is as follows:

[0877] 1. Receive audio or image data transmitted from the smart glasses.

[0878] 2. Preprocess the received data to remove noise, adjust the volume, remove background, and correct color.

[0879] 3. The pre-processed data is fed into a generative AI model for technical analysis.

[0880] 4. Analyze the emotional state of the worker using an emotion engine.

[0881] 5. Generate specific improvements based on these results.

[0882] 6. Generate feedback based on improvements and emotions and display it on the smart glasses display.

[0883] Device Role

[0884] The device (smart glasses) captures and transmits data using the following hardware and software:

[0885] Hardware:

[0886] Camera: Capture footage of your work

[0887] Microphone: Captures audio

[0888] Display: Show feedback

[0889] software:

[0890] Real-time capture and data transmission applications

[0891] The process on the terminal is as follows:

[0892] 1. Use a camera and microphone to capture work video and audio in real time.

[0893] 2. Upload the captured data to the server using the HTTPS protocol.

[0894] User Roles

[0895] The user (worker) wears the smart glasses and performs the work. The user's operation procedure is as follows.

[0896] 1. Put on the smart glasses and start working.

[0897] 2. The camera and microphone in the smart glasses capture the work situation and send it to the server.

[0898] 3. The feedback sent back from the server is viewed on the smart glasses display.

[0899] 4. Correct and improve your work based on feedback.

[0900] Specific examples

[0901] Example: When a factory worker assembles parts, the camera in the smart glasses captures the work process and the microphone records the work instructions. The data is sent to a server, and feedback such as "Part A is facing the wrong way" or "Please work more relaxedly" is displayed in real time based on technical evaluation and emotional state analysis.

[0902] Prompt Sentence Examples

[0903] Image data: Video of parts assembly

[0904] Audio data: "Place part A on part B."

[0905] Issue: Offer advice on how to improve their work and their emotions.

[0906] In this way, users can receive feedback based on their work evaluation and emotions in real time, allowing them to improve their skills efficiently.

[0907] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0908] Step 1:

[0909] The user puts on the smart glasses and starts working. The smart glasses capture the work video with a camera and the audio with a microphone. These data (input) are used to record the work progress and audio instructions in real time. The real-time captured video data and audio data are generated as output.

[0910] Step 2:

[0911] The device (smart glasses) securely uploads the captured video and audio data to the server using the HTTPS protocol. This process ensures the data is securely transmitted to the server. The input is the captured video and audio data, and the output is the video and audio data transmitted to the server.

[0912] Step 3:

[0913] The server preprocesses the received video and audio data. Specifically, it performs noise reduction and volume adjustment on the audio data, and background removal and color correction on the video data. The input is the video and audio data sent to the server, and the output is the preprocessed data. This is done using libraries such as OpenCV and SciPy.

[0914] Step 4:

[0915] The server inputs the preprocessed data into the generative AI model and performs technical analysis. The generative AI model analyzes the input data based on the data it has previously learned. The input is the preprocessed data, and the output is the analysis results (specific technical improvements).

[0916] Step 5:

[0917] The server uses an emotion engine to analyze the user's emotional state from voice tone and visual expressions. The input is preprocessed data, and the output is the recognition result of the emotional state. This is done based on voice tone, patterns, visual expressions, etc.

[0918] Step 6:

[0919] The server generates specific improvement points and feedback according to emotions based on the results of technical analysis and emotion recognition. The inputs are the analysis results and emotion recognition results, and the output is feedback to be provided to the user (advice on how to improve work and respond to emotions).

[0920] Step 7:

[0921] The server sends the generated feedback to the device (smart glasses) using the HTTPS protocol, so that the user can check the feedback in real time. The input is the feedback data, and the output is the feedback sent to the smart glasses.

[0922] Step 8:

[0923] The terminal (smart glasses) displays the received feedback on its display. The user checks the feedback in real time and corrects and improves their work. The input is the feedback sent from the server, and the output is the feedback displayed on the smart glasses display.

[0924] In this way, data is captured, transferred, processed, and feedback is provided at each step, allowing users to receive real-time technical improvement and emotionally sensitive guidance.

[0925] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0926] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0927] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0928] [Fourth embodiment]

[0929] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0930] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0931] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0932] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0933] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0934] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0935] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0936] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0937] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0938] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0939] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0940] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0941] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0942] The present invention relates to a technology improvement support system using voice data and image data. How the present invention is put into practice will be described below in detail.

[0943] First, the user launches a dedicated application on their smartphone or PC. Through the application, the user creates audio data of their own performance or image data of a painting they have drawn. For example, if the user is singing, they record the audio using the smartphone's microphone. If they are painting, they use the smartphone's camera to take a photo of the work.

[0944] The device then uploads the audio or image data to a server via the application. The HTTPS protocol is used for this upload, ensuring secure transmission of the data to the server. The received data is then tagged with the user's identification information, ensuring accurate feedback to the user.

[0945] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0946] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0947] Once the generative AI model has completed its analysis, the server generates specific improvements based on the results. For example, feedback such as "certain notes are off" or "the rhythm is inconsistent" or "the color scheme is biased" or "the composition is unstable" may be included. The server also provides this feedback in written form, making it easy for users to understand. At the same time, if the data is audio, a corrected version of the audio is also generated.

[0948] Finally, the server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed within the application. The user can learn how to improve their skills based on the specific feedback and use it for their next practice.

[0949] In addition, the server will partner with relevant professional and educational institutions to share data and retrain the AI ​​model to improve its accuracy and increase feedback variations, thereby continuously improving the system's effectiveness and reliability and providing a higher level of support to users.

[0950] The above is an embodiment of the present invention, showing the specific operation of a system for supporting individual skill improvement using audio data and image data.

[0951] The processing flow will be explained below.

[0952] Step 1: Enter your data

[0953] The user starts a dedicated application using a smartphone or a PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image).

[0954] Step 2: Upload your data

[0955] The device uploads the recorded or captured audio and image data to a server, using the HTTPS protocol for secure communication.

[0956] Step 3: Receiving the data

[0957] The server receives the voice data and image data sent from the terminal, as well as the user's identification information.

[0958] Step 4: Preprocessing the audio data

[0959] The server performs noise reduction and volume adjustment on the received audio data, using FFT (Fast Fourier Transform) and noise reduction algorithms.

[0960] Step 5: Preprocessing the image data

[0961] The server performs background removal and color correction on the received image data, using image processing libraries such as OpenCV.

[0962] Step 6: Analysis by generative AI model

[0963] The server inputs preprocessed audio or image data into the generative AI model and analyzes the data. For audio data, it evaluates pitch, rhythm, and vocal accuracy, while for image data, it evaluates composition, color usage, and texture.

[0964] Step 7: Generate improvements

[0965] The server generates specific improvements based on the analysis results obtained from the generative AI model, such as noting that certain notes are off, the rhythm is inconsistent, the color scheme is biased, or the composition is unstable.

[0966] Step 8: Generate corrected audio and video

[0967] The server corrects the audio data and image data as necessary and generates corrected data. In the case of audio data, it creates an audio file with corrected pitch and rhythm.

[0968] Step 9: Provide feedback

[0969] The server sends the generated feedback and correction data back to the user, using notifications to allow the user to view the feedback within the application.

[0970] Step 10: Review feedback

[0971] Users can review the feedback provided within the application, understand specific areas for improvement, and apply the corrective data to their next practice.

[0972] Step 11: Business partnerships and data utilization

[0973] The server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources to retrain the generative AI model and improve the accuracy of the system.

[0974] The above is a description of the specific operations for each processing step. Through this series of processes, users can improve their skills efficiently and effectively.

[0975] Example 1

[0976] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0977] Conventional systems that support individual skill improvement using voice and image data suffer from insufficient preprocessing and analysis, resulting in poor quality feedback provided to users and making it difficult to identify specific areas for improvement. There are also security concerns regarding data transmission. Furthermore, because analysis requires specialized knowledge, users often fail to achieve the results they intended. There is a need for a system that can solve these issues and enable users to truly experience improvements in their skills.

[0978] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0979] In this invention, the server includes: a means for a user to create audio data or image data; a means for securely transmitting the audio data or image data; a means for preprocessing the audio data, such as noise reduction and volume adjustment, or background removal and color correction for the image data; a means for analyzing the preprocessed data using a generative AI model; a means for generating specific improvements based on the analysis results; and a means for providing the improvements to the user. This allows users to receive high-quality, specific feedback, making it easier for them to realize their technical improvements. Furthermore, secure data transmission also improves reliability in terms of security.

[0980] "Means for users to create audio data or image data" refers to a function that allows users to record or take pictures using a dedicated application.

[0981] The "means for securely transmitting the audio data or image data" refers to a function for encrypting data using a security protocol (for example, HTTPS) and securely transmitting the data to a server.

[0982] "Means for performing pre-processing such as noise removal and volume adjustment for audio data, or background removal and color correction for image data" refers to a function that performs processing to improve the quality of data using algorithms and libraries such as FFT for audio data and OpenCV for image data.

[0983] "Means for analyzing preprocessed data using a generative AI model" refers to the function of evaluating preprocessed audio data or image data using an AI model that has previously trained on expert feedback data.

[0984] "Means for generating specific improvements based on the analysis results" refers to a function that generates specific improvements that the user should make (for example, pitch discrepancies or unstable composition) based on the analysis results of the generation AI model.

[0985] "Means for providing the user with the improvements" refers to a function for returning specific improvements obtained as a result of the analysis to the user as text or correction data and notifying them via the application.

[0986] The present invention relates to a technology improvement support system using voice data and image data. How the present invention is put into practice will be described below in detail.

[0987] First, the user launches a dedicated application on their smartphone or PC. Through the application, the user can record audio data of their own performance or create image data of a painting they have drawn. For example, if the user is singing, they can record the audio using the smartphone's microphone. If they are painting, they can take a photo of the work using the smartphone's camera.

[0988] The device then uploads the audio or image data to a server via the application. The HTTPS protocol is used for this upload, ensuring secure transmission of the data to the server. The received data is then tagged with the user's identification information, ensuring accurate feedback to the user.

[0989] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[0990] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[0991] Once the analysis is complete, the server generates specific improvements based on the results. For example, feedback such as "certain notes are off" or "the rhythm is inconsistent" or "the color scheme is biased" or "the composition is unstable" may be included. The server also provides this feedback in written form, making it easy for users to understand. At the same time, if the data is audio, a corrected version of the audio is also generated.

[0992] Finally, the server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed within the application. The user can learn how to improve their skills based on the specific feedback and use it for their next practice.

[0993] In addition, the server will partner with relevant professional and educational institutions to share data and retrain the AI ​​model to improve its accuracy and increase feedback variations, thereby continuously improving the system's effectiveness and reliability and providing a higher level of support to users.

[0994] Examples of specific examples and prompts

[0995] Example 1: Analyzing audio data

[0996] Users record their own singing on their smartphones and upload the audio data to a server via the application.

[0997] The device securely transmits the audio data to the server using the HTTPS protocol.

[0998] The server uses FFT to remove noise and adjust the volume, while the generative AI model analyzes pitch and rhythm and generates feedback that a particular note is out of tune.

[0999] The user will receive feedback from the application that indicates that a particular note is off, and can use this feedback to improve their next practice.

[1000] Example prompt:

[1001] Analyze my singing performance. I provide audio data. You rate this audio for pitch and rhythmic accuracy and suggest areas for improvement.

[1002] Example 2: Image data analysis

[1003] The user takes a photo of the painting they have created with their smartphone and uploads the image data to the server via the application.

[1004] The device securely transmits the image data to the server using the HTTPS protocol.

[1005] The server uses OpenCV to remove the background of the image and correct the color tone, and the generative AI model analyzes the composition and color usage and generates feedback such as "the color scheme is biased."

[1006] The user can check the feedback in the application that the color scheme is biased and use it to create their next work.

[1007] Example prompt:

[1008] Please analyze a painting. Provide image data. Evaluate the composition and color usage of the work, and suggest areas for improvement.

[1009] The above is an embodiment of the present invention, showing the specific operation of a system for supporting users in improving their individual skills using audio data and image data.

[1010] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1011] Step 1:

[1012] A user launches a dedicated application using a smartphone or PC. The user creates the audio data they want to record or the image data they want to capture. For example, if the user wants to sing, they can use the application's recording function to record the audio. For image data, they can use the application's camera function to take a photo of a painting. The input data can be audio or image.

[1013] Step 2:

[1014] After the user finishes recording or shooting, they check the data in the application and press the save button. This saves the audio or image data to the device. The input data is the audio or image that the user checked and edited, and the saved audio or image file is obtained as output data.

[1015] Step 3:

[1016] The device uploads the saved audio and image data to the server using the HTTPS protocol. The user's identification information is also sent, so accurate feedback can be returned to the user. The input data is the audio file, image file, and user identification information, and is uploaded to the server based on this information.

[1017] Step 4:

[1018] The server preprocesses the received audio and image data. For audio data, it uses FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, it uses OpenCV to remove the background and correct the color tone. The input data is an audio file or an image file, and based on this, data processing such as noise removal, volume adjustment, background removal, and color correction is performed. The output data is a preprocessed audio file or image file.

[1019] Step 5:

[1020] The preprocessed data is input into a generative AI model stored on the server. The generative AI model has previously studied feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, and in the case of image data, it evaluates composition, color usage, and texture. The input data are preprocessed audio and image files, and the output data are the analysis results.

[1021] Step 6:

[1022] The server generates specific improvements based on the analysis results of the generative AI model. For example, it evaluates audio such as "certain pitches are off" or "rhythm is inconsistent," and images such as "color scheme is biased" or "composition is unstable." The input data is the analysis results, and the output data is generated in document format as specific improvements.

[1023] Step 7:

[1024] The server sends the generated feedback information back to the user. The feedback is notified via the application and can be viewed by the user within the application. The input data is the generated feedback information, which is used to notify the user. The output data is the specific feedback the user receives.

[1025] Step 8:

[1026] Users can check the provided feedback and use it to improve their skills. They can use the feedback to improve their next practice or work and improve their skills. The input data is feedback information, and the output data is specific actions that will lead to the user's skill improvement.

[1027] (Application example 1)

[1028] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1029] Conventional technology improvement support systems have problems with the insufficient accuracy and speed of analysis and feedback of voice or image data. Another issue is that the devices users can use are limited, making it difficult to provide convenience across a variety of devices. Furthermore, there are insufficient means for providing specific improvements based on the analysis results, and users often lack guidance on how to actually implement improvements.

[1030] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1031] In this invention, the server includes means for receiving voice data or image data, means for preprocessing the voice data or image data, means for analyzing the preprocessed data using a generative AI model, means for generating specific improvements based on the analysis results, means for providing the improvements to the user, means for uploading data to a cloud server and receiving feedback, and means for installation on a smartphone or head-mounted display. This enables highly accurate and rapid analysis and feedback of improvements. Furthermore, the system allows users to use the system on a variety of devices, greatly improving convenience. Furthermore, providing specific improvements allows users to obtain useful guidance for actually improving their performance.

[1032] definition statement

[1033] "Audio data" refers to recorded audio stored in digital format.

[1034] "Image data" refers to data that is a captured or generated image stored in digital format.

[1035] "Preprocessing" refers to initial data processing such as noise removal and color correction to improve the quality of audio and image data.

[1036] A "generative AI model" is an artificial intelligence model that learns using large amounts of data and analyzes and evaluates audio and images.

[1037] A "cloud server" is a remote server that stores, analyzes, and provides data over the Internet.

[1038] A "smartphone" is a portable information terminal with advanced computing power and multiple functions.

[1039] A "head-mounted display" is a device that is worn on the head and displays a display directly within the field of view.

[1040] "User" means a person or entity that uses the system to upload audio or image data and receive feedback.

[1041] "Preprocessed data" refers to audio data or image data that has undergone preprocessing such as noise removal and color correction.

[1042] "Specific improvements" are clear suggestions and advice for improving the user's performance based on the analysis of the generative AI model.

[1043] MODE FOR CARRYING OUT THE INVENTION

[1044] This invention relates to a technology improvement support system using voice and image data. This system receives voice or image data, preprocesses them, analyzes them using a generative AI model, and provides specific feedback on improvements to the user. The following describes how this system is implemented.

[1045] First, the user launches a dedicated application using a smartphone or head-mounted display. The user then creates audio data by recording their own performance or image data by photographing it. For example, if the user is singing, they can record the audio using the smartphone's microphone, or if they are painting, they can take a photo of the artwork using the smartphone's camera.

[1046] The device then uploads the audio or image data to a cloud server via the application. The HTTPS protocol is used for uploading, and the data is securely sent to the server. The server then adds user identification information to the received data, ensuring appropriate feedback.

[1047] The server preprocesses the received audio and image data. For audio data, the server uses FFT (Fast Fourier Transform) and filtering algorithms to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and correct color. This preprocessing ensures that subsequent analysis is performed with high accuracy.

[1048] The preprocessed data is input into a generative AI model stored on the server. The generative AI model is trained based on feedback data collected from music and art experts, enabling highly accurate analysis. Audio data is evaluated for pitch, rhythm, and vocal accuracy, while image data is evaluated for composition, color usage, and texture.

[1049] Once the analysis is complete, the server generates specific improvements based on the results. These improvements include feedback such as "certain notes are off" or "the rhythm is inconsistent," as well as "the color scheme is biased" or "the composition is unstable." Corrected audio and images may also be generated.

[1050] Finally, the server sends the generated feedback information back to the user. The feedback is notified to the user through the application and can be viewed within the application. The user can learn strategies to improve their skills based on this feedback and use it for their next practice.

[1051] As a specific example, analysis is performed by inputting the following prompt sentence into the generative AI model.

[1052] Example of prompt for audio data:

[1053] Please analyze the following audio data and indicate areas for improvement in pitch and rhythm:

[1054] [Audio data]

[1055] Example prompt for image data:

[1056] Analyze the image below and suggest improvements to the color scheme and composition:

[1057] [Image data]

[1058] This system enables accurate and rapid analysis and feedback on improvements, allows users to use the system across a variety of devices, and provides specific improvements, providing useful guidance for users to actually improve their performance.

[1059] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1060] System processing steps

[1061] Step 1:

[1062] The user launches a dedicated application using a smartphone or head-mounted display. The user then creates audio data by recording their own performance or image data by taking photographs. The input is raw audio or image data, obtained by recording or taking photographs. The output is audio or image data stored within the device.

[1063] Step 2:

[1064] The device uploads audio or image data to the cloud server through the application. The HTTPS protocol is used during uploading to ensure data security. The input is the audio or image data stored in the device, and the output is the data uploaded to the cloud server.

[1065] Step 3:

[1066] The server preprocesses the received data. For audio data, FFT (Fast Fourier Transform) and filtering algorithms are used to remove noise and adjust the volume. For image data, OpenCV is used to remove background and correct color. The input is raw data uploaded to the cloud server, and the output is preprocessed data.

[1067] Step 4:

[1068] The server inputs the preprocessed data into the generative AI model. The generative AI model is trained based on feedback data collected in advance from experts, enabling highly accurate analysis. The input is preprocessed data, and the output is the analysis results. In the case of audio data, pitch, rhythm, and vocal accuracy are evaluated, while in the case of image data, composition, color usage, and texture are evaluated.

[1069] Step 5:

[1070] The server generates specific improvements based on the analysis results. This includes feedback such as "certain pitches are off" or "the rhythm is inconsistent," as well as "the color scheme is biased" or "the composition is unstable." The input is the analysis results obtained from the generative AI model, and the output is specific feedback on improvements.

[1071] Step 6:

[1072] The server sends the generated feedback information back to the user. The feedback is notified through the application and can be viewed by the user within the application. The input is specific feedback for improvement, and the output is notification and display to the user.

[1073] This system allows users to learn strategies for improving their own performance based on specific areas for improvement and apply them to their next practice.In addition, the system enables highly accurate and rapid analysis, and can be used on a variety of devices, greatly improving convenience.

[1074] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1075] This invention combines an emotion engine with a technology improvement support system that uses voice data and image data. How this invention is put into practice will be described below in detail.

[1076] First, a user launches a dedicated application on a smartphone or PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image). For example, if a user sings, they use the smartphone's microphone to record the audio. If it is a painting, they use the smartphone's camera to take a photo of the work.

[1077] The device then uploads the audio or image data to a server, using the HTTPS protocol for secure communication. The received data is tagged with the user's identification information, ensuring accurate feedback to the user.

[1078] The server preprocesses the received audio or image data. For audio data, the server uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For image data, the server uses image processing libraries such as OpenCV to remove background and perform color correction. This preprocessing makes subsequent analysis more accurate.

[1079] The preprocessed data is then input into a generative AI model on the server. This model has previously learned from feedback data from music and art experts, enabling highly accurate analysis. In the case of audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy, while in the case of image data, it evaluates composition, color usage, and texture.

[1080] Once the generative AI model has completed its analysis, the server uses an emotion engine to recognize the user's emotions. The server analyzes the voice tone and patterns from the voice data to identify the user's emotional state (e.g., joy, sadness, anger, etc.). The server also analyzes facial expressions and composition from the image data to identify the user's emotional state.

[1081] The server generates specific improvements based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, in addition to pointing out issues such as "certain pitches are off," "rhythms are inconsistent," "color schemes are biased," and "composition is unstable," the server also provides advice based on the user's emotional state. For example, if the emotion recognition result is "nervous," specific advice such as "how to speak in a relaxed manner" is provided.

[1082] The server then corrects the audio data and image data as necessary and generates corrected data. In the case of audio data, it creates an audio file with corrected pitch and rhythm.

[1083] Finally, the server sends the generated feedback and correction data back to the user. The feedback and correction data are notified via the application and can be viewed within the application. The user can understand specific improvements and emotion-based advice and use it for their next practice.

[1084] Additionally, the server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources to retrain the generative AI model and emotion engine, thereby improving the accuracy of the system and providing a higher level of support to users.

[1085] The above describes the mode for carrying out the present invention, and shows the specific operation of a system for supporting individual skill improvement using voice data and image data. By combining it with an emotion engine, feedback is provided that takes into account the user's emotional state, and more effective skill improvement can be expected.

[1086] The processing flow will be explained below.

[1087] Step 1: Enter your data

[1088] A user uses a smartphone or PC to launch a dedicated application and record or photograph audio data (e.g., singing voice) or image data (e.g., painting image). For example, when recording singing voice, the user uses the smartphone's microphone and presses the recording start button within the application. In the case of a painting, the user takes a photo of the work using the smartphone's camera.

[1089] Step 2: Upload your data

[1090] The device uploads recorded or captured audio and image data to a server via the application. Secure communication is performed using the HTTPS protocol. When uploading data, user identification information is also sent.

[1091] Step 3: Receiving the data

[1092] The server receives the voice data or image data sent from the terminal, and then classifies and stores the data appropriately based on the user identification information.

[1093] Step 4: Preprocessing the audio data

[1094] The server performs noise reduction and volume adjustment on the received audio data. Specifically, it applies a noise reduction algorithm using FFT (Fast Fourier Transform) to remove background noise. If the volume is not consistent, it also applies a volume adjustment algorithm.

[1095] Step 5: Preprocessing the image data

[1096] The server performs background removal and color correction on the received image data. For background removal, it uses an image processing library such as OpenCV, and for color correction, it uses an automatic white balance adjustment algorithm.

[1097] Step 6: Analysis by generative AI model

[1098] The server inputs preprocessed audio or image data into the generative AI model, which analyzes the data. For audio data, the AI ​​model evaluates pitch, rhythm, and vocal accuracy. For image data, the AI ​​model evaluates composition, color, and texture.

[1099] Step 7: Emotion Recognition with the Emotion Engine

[1100] After obtaining the analysis results from the generative AI model, the server uses an emotion engine to recognize the user's emotions. For voice data, it analyzes the voice tone and patterns to identify the user's emotional state (e.g., joy, sadness, anger, etc.). For image data, it uses facial expression analysis algorithms and composition analysis to identify the emotional state.

[1101] Step 8: Generate recommendations based on improvements and sentiment

[1102] The server generates specific improvements based on the analysis results and emotion recognition results obtained from the generative AI model. For example, it points out issues such as "certain pitches are off," "rhythms are inconsistent," "color schemes are biased," and "composition is unstable," and provides advice based on the emotional state (e.g., "how to speak in a relaxed manner").

[1103] Step 9: Generate corrected audio and images

[1104] The server generates modified audio and images based on the original audio or image data. In the case of audio data, it creates an audio file with corrected pitch and rhythm. In the case of image data, it creates an image file with corrected composition and color tone.

[1105] Step 10: Provide feedback

[1106] The server returns the generated feedback information and correction data to the user, who is notified of the feedback and correction data using a notification function so that the user can check it within the application.

[1107] Step 11: Review feedback

[1108] Users can review the feedback and correction data generated within the application, learn specific improvements and emotion-based advice to improve their next practice.

[1109] Step 12: Business partnerships and data utilization

[1110] The server will partner with stakeholders such as music schools and museums to acquire new datasets and feedback resources, and retrain the generative AI model and emotion engine to improve the accuracy of the system, thereby providing a higher level of support to users.

[1111] This concludes the explanation of the specific operations for each processing step. This series of processes allows users to efficiently and effectively improve their skills, and by adding emotion recognition, they can receive more personalized feedback.

[1112] Example 2

[1113] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1114] Conventional skill improvement support systems do not take into account the user's emotional state when analyzing voice and image data, resulting in feedback that is not adapted to the user's psychological state. Furthermore, the feedback provided is often vague, making it unclear how users should use it to improve their skills. Furthermore, data correction is often done manually, resulting in low efficiency.

[1115] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1116] In this invention, the server includes means for receiving voice data or image data, means for preprocessing the voice data or image data, means for analyzing the preprocessed data using a generative AI model, means for generating specific improvements based on the analysis results and emotion recognition results from an emotion engine, means for providing the user with advice corresponding to the improvements and the user's emotional state, means for correcting the data based on the provided improvements and advice, and means for returning the corrected data and feedback information to the user. This not only supports the user's technical improvement, but also enables highly accurate feedback that takes into account the user's emotional state at the time. Furthermore, the automatic data correction function enables efficient technical improvement support.

[1117] "Audio data" means sound signals recorded using a microphone or other sound recording device.

[1118] "Image data" means signals containing visual information captured using a camera or other image capturing device.

[1119] "Preprocessing" refers to processing carried out to improve the quality of data before data analysis, and includes noise removal and color correction.

[1120] A "generative AI model" refers to an artificial intelligence algorithm that is trained in advance on a large dataset and analyzes and evaluates audio and image data.

[1121] "Emotion Engine" means a software or hardware component for recognizing a user's emotional state from audio and / or visual data.

[1122] "Specific improvements" refers to specific changes or corrections that users should make to improve the technology, based on the analysis results of the generative AI model and the recognition results of the emotion engine.

[1123] "Advice" means guidance or advice to a User based on the results of the generative AI model and emotion engine.

[1124] "Data Correction" means any changes or improvements to the data based on analysis and feedback.

[1125] "Feedback Information" means information, including analytical results and improvements derived from the results of the Generative AI Model and Emotion Engine.

[1126] This invention combines an emotion engine with a technology improvement support system that uses voice data and image data. How this invention is put into practice will now be described in detail.

[1127] A user uses a smartphone or PC to launch a dedicated application and record or photograph audio data (e.g., singing voice) or image data (e.g., painting image). For example, if a user is singing, they can record "Happy Birthday" using the smartphone's microphone. If a user is taking a picture of a painting, they can use the smartphone's camera to photograph a landscape.

[1128] The device uploads the collected audio or image data to a server. This is done securely using the HTTPS protocol. The received data is tagged with the user's identification information, ensuring accurate feedback is sent back to the user.

[1129] The server preprocesses the received audio or image data. For audio data, it uses algorithms such as FFT (Fast Fourier Transform) to remove noise and adjust the volume. For example, it filters out background noise and equalizes the overall volume. For image data, it uses image processing libraries such as OpenCV to remove the background and perform color correction. For example, it converts the background of an image to black and white and extracts only the main parts.

[1130] The preprocessed data is input into a generative AI model stored on the server. This generative AI model has previously learned from feedback data from music and art experts, allowing it to perform highly accurate analysis. In the case of audio data, the generative AI model evaluates the accuracy of pitch, rhythm, and vocalization. For example, it may evaluate the data as "the C4 note is low" or "the rhythm is unstable." In the case of image data, it evaluates the composition, color usage, and texture. For example, it may evaluate the data as "the color scheme is unbalanced" or "the composition is biased."

[1131] Once the analysis by the generative AI model is complete, the server uses an emotion engine to analyze the user's emotions. The tone and patterns of the voice data are analyzed to identify the user's emotional state (e.g., joy, tension, relief, etc.). The facial expressions and composition of the image data are also analyzed to similarly identify the emotional state. For example, if the voice tone is bright and tense, it is judged to be "joy," while if the voice is low and the tone is unstable, it is judged to be "tension."

[1132] The server generates specific areas for improvement and advice based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, specific suggestions such as "Your C4 note is low, so you need to practice producing an A4 note accurately" and emotion-based advice such as "If you're nervous, practice vocal exercises that incorporate deep breathing to relax" are provided.

[1133] Furthermore, the server corrects the audio and image data as needed. In the case of audio data, it generates an audio file with corrected pitch and rhythm. For example, it corrects the pitch and rhythm of the singing data of "Happy Birthday" to bring it closer to the correct pitch and rhythm. In the case of image data, it readjusts the color tone and composition.

[1134] Finally, the server sends the generated feedback information and correction data back to the user, who can then review the provided feedback and correction data through a dedicated application, allowing the user to understand specific improvements and emotion-based advice and use it to improve their skills.

[1135] An example of a prompt sentence for voice data (singing voice):

[1136] "Please analyze my recording of "Happy Birthday" and let me know how I can improve it."

[1137] For image data (pictorial images):

[1138] "Please give me feedback on the composition and color scheme of this landscape painting."

[1139] In this way, users can receive feedback and advice on specific techniques to improve their skills and use it in their practice and creative work.

[1140] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1141] Step 1:

[1142] The user starts a dedicated application on their smartphone or PC and records or photographs audio data (e.g., singing voice) or image data (e.g., painting image). Specifically, the user uses the smartphone's microphone to sing "Happy Birthday" and record the audio. This generates input audio data or image data, which then becomes the input for the next step.

[1143] Step 2:

[1144] The device uploads the collected voice or image data to a server. Secure communication is performed using the HTTPS protocol. The data is accompanied by the user's identification information, ensuring accurate feedback to the user. The input is voice or image data, and the output is the data transferred to the server.

[1145] Step 3:

[1146] The server preprocesses the received audio or image data. For audio data, it uses FFT (Fast Fourier Transform) to remove noise and adjust the volume. Specifically, it filters out background noise and keeps the volume constant. For image data, it uses an image processing library such as OpenCV to remove the background and correct the color tone. The preprocessed data is the input data, and the results of this processing are input to the next step.

[1147] Step 4:

[1148] The server inputs preprocessed audio or image data into the generative AI model. The generative AI model has previously studied expert feedback data, allowing it to perform highly accurate analysis. In the case of audio data, the generative AI model evaluates pitch, rhythm, and accuracy of pronunciation. Specifically, it evaluates whether "the C4 note is low" or "the rhythm is unstable." In the case of image data, it evaluates composition, color usage, and texture. This becomes the analyzed data output.

[1149] Step 5:

[1150] After the server has completed data analysis using the generative AI model, it uses an emotion engine to analyze the user's emotions. In the case of voice data, it analyzes the voice tone and patterns to identify the user's emotional state. For example, it may determine that "if the voice tone is bright and high-tension, it is 'joy'" or "if the voice tone is low and unstable, it is 'tension'." In the case of image data, it analyzes facial expressions and composition to similarly recognize the emotional state. This results in the user's emotion recognition results being output.

[1151] Step 6:

[1152] The server generates specific areas for improvement and advice based on the analysis results obtained from the generative AI model and the emotion recognition results from the emotion engine. For example, technical advice such as "Your C4 note is low, so you need to practice producing an A4 accurately" and emotional advice such as "If you're nervous, practice vocalization by incorporating deep breathing to relax" are generated. These are output as areas for improvement and advice.

[1153] Step 7:

[1154] The server corrects the audio data and image data as necessary. In the case of audio data, it generates an audio file with corrected pitch and rhythm. For example, it corrects the pitch and rhythm of the singing data of "Happy Birthday" to bring it closer to the correct pitch and rhythm. In the case of image data, it readjusts the color tone and composition. This is output as corrected data.

[1155] Step 8:

[1156] The server sends the generated feedback information and correction data back to the user, who can then review the provided feedback and correction data through a dedicated application. This allows the user to understand specific improvements and emotional advice, which can be used to improve their skills. This is the final output.

[1157] (Application example 2)

[1158] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1159] In conventional skill improvement support systems, it was difficult to simultaneously provide feedback that took into account the user's emotional state along with specific points for improvement in the user's skills. Furthermore, while real-time technical guidance is required, particularly in actual workplaces such as factories, there was a lack of a means to provide immediate feedback and support the improvement of workers' skills.

[1160] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving audio data or image data, means for preprocessing the audio data or the image data, and means for analyzing the preprocessed data using a generative AI model. This includes means for capturing work video and audio in real time using the camera and microphone of the smart glasses and uploading the data to the server, means for generating feedback in the server based on work improvement points and emotions and displaying it on the display of the smart glasses, means for generating specific improvement points based on the analysis results, and means for providing the improvement points to the user. This makes it possible to provide specific feedback in real time based on the user's technical improvement points and emotions.

[1161] "Audio Data" means information for recording, transmitting, and analyzing audio in digital or analog form.

[1162] "Image data" refers to information for recording, transmitting, and analyzing visual information in digital or analog format.

[1163] "Preprocessing" refers to the initial processing steps taken to convert data into a form suitable for analysis.

[1164] A "generative AI model" is an artificial intelligence system that uses pre-trained data to evaluate and generate new data.

[1165] The "emotion engine" is a system for analyzing and identifying a user's emotional state from voice and image data.

[1166] "Real-time capture" means recording and processing current events and actions as they occur.

[1167] "Smart glasses" are wearable devices that incorporate sensors such as a camera and microphone and provide visual information to the user.

[1168] "Feedback" refers to suggestions and advice for improvement provided based on analysis results and evaluations.

[1169] "Server" means a computer system established for the purpose of processing, storing, and managing data.

[1170] This invention is a system that provides real-time technical guidance to factory workers while they are working using smart glasses. The roles of the server, terminal, and user are as follows:

[1171] Server Roles

[1172] The server uses the following hardware and software to process and analyze data:

[1173] Hardware: A server with a powerful processor and plenty of memory

[1174] software:

[1175] OpenCV: Image processing library

[1176] SciPy: an audio processing library

[1177] Transformers: A library of generative AI models for sentiment analysis

[1178] HTTPS protocol: secure data communication

[1179] The data processing flow is as follows:

[1180] 1. Receive audio or image data transmitted from the smart glasses.

[1181] 2. Preprocess the received data to remove noise, adjust the volume, remove background, and correct color.

[1182] 3. The pre-processed data is fed into a generative AI model for technical analysis.

[1183] 4. Analyze the emotional state of the worker using an emotion engine.

[1184] 5. Generate specific improvements based on these results.

[1185] 6. Generate feedback based on improvements and emotions and display it on the smart glasses display.

[1186] Device Role

[1187] The device (smart glasses) captures and transmits data using the following hardware and software:

[1188] Hardware:

[1189] Camera: Capture footage of your work

[1190] Microphone: Captures audio

[1191] Display: Show feedback

[1192] software:

[1193] Real-time capture and data transmission applications

[1194] The process on the terminal is as follows:

[1195] 1. Use a camera and microphone to capture work video and audio in real time.

[1196] 2. Upload the captured data to the server using the HTTPS protocol.

[1197] User Roles

[1198] The user (worker) wears the smart glasses and performs the work. The user's operation procedure is as follows.

[1199] 1. Put on the smart glasses and start working.

[1200] 2. The camera and microphone in the smart glasses capture the work situation and send it to the server.

[1201] 3. The feedback sent back from the server is viewed on the smart glasses display.

[1202] 4. Correct and improve your work based on feedback.

[1203] Specific examples

[1204] Example: When a factory worker assembles parts, the camera in the smart glasses captures the work process and the microphone records the work instructions. The data is sent to a server, and feedback such as "Part A is facing the wrong way" or "Please work more relaxedly" is displayed in real time based on technical evaluation and emotional state analysis.

[1205] Prompt Sentence Examples

[1206] Image data: Video of parts assembly

[1207] Audio data: "Place part A on part B."

[1208] Issue: Offer advice on how to improve their work and their emotions.

[1209] In this way, users can receive feedback based on their work evaluation and emotions in real time, allowing them to improve their skills efficiently.

[1210] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1211] Step 1:

[1212] The user puts on the smart glasses and starts working. The smart glasses capture the work video with a camera and the audio with a microphone. These data (input) are used to record the work progress and audio instructions in real time. The real-time captured video data and audio data are generated as output.

[1213] Step 2:

[1214] The device (smart glasses) securely uploads the captured video and audio data to the server using the HTTPS protocol. This process ensures the data is securely transmitted to the server. The input is the captured video and audio data, and the output is the video and audio data transmitted to the server.

[1215] Step 3:

[1216] The server preprocesses the received video and audio data. Specifically, it performs noise reduction and volume adjustment on the audio data, and background removal and color correction on the video data. The input is the video and audio data sent to the server, and the output is the preprocessed data. This is done using libraries such as OpenCV and SciPy.

[1217] Step 4:

[1218] The server inputs the preprocessed data into the generative AI model and performs technical analysis. The generative AI model analyzes the input data based on the data it has previously learned. The input is the preprocessed data, and the output is the analysis results (specific technical improvements).

[1219] Step 5:

[1220] The server uses an emotion engine to analyze the user's emotional state from voice tone and visual expressions. The input is preprocessed data, and the output is the recognition result of the emotional state. This is done based on voice tone, patterns, visual expressions, etc.

[1221] Step 6:

[1222] The server generates specific improvement points and feedback according to emotions based on the results of technical analysis and emotion recognition. The inputs are the analysis results and emotion recognition results, and the output is feedback to be provided to the user (advice on how to improve work and respond to emotions).

[1223] Step 7:

[1224] The server sends the generated feedback to the device (smart glasses) using the HTTPS protocol, so that the user can check the feedback in real time. The input is the feedback data, and the output is the feedback sent to the smart glasses.

[1225] Step 8:

[1226] The terminal (smart glasses) displays the received feedback on its display. The user checks the feedback in real time and corrects and improves their work. The input is the feedback sent from the server, and the output is the feedback displayed on the smart glasses display.

[1227] In this way, data is captured, transferred, processed, and feedback is provided at each step, allowing users to receive real-time technical improvement and emotionally sensitive guidance.

[1228] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1229] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1230] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1231] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1232] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1233] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1234] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1235] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1236] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1237] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1238] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1239] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1240] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1241] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1242] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1243] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1244] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1245] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1246] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1247] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1248] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1249] The following is further disclosed regarding the above embodiment.

[1250] (Claim 1)

[1251] means for receiving audio data or image data;

[1252] means for preprocessing the audio data or the image data;

[1253] A means of analyzing the pre-processed data using a generative AI model; and

[1254] means for generating specific improvements based on the analysis results;

[1255] means for providing said improvements to a user;

[1256] A technology improvement support system including:

[1257] (Claim 2)

[1258] 10. The skill improvement support system of claim 1, further comprising means for performing noise removal and volume adjustment in preprocessing of the audio data.

[1259] (Claim 3)

[1260] 10. The technology improvement support system of claim 1, further comprising means for performing background removal and color correction in preprocessing of image data.

[1261] "Example 1"

[1262] (Claim 1)

[1263] means for a user to create audio data or image data;

[1264] means for securely transmitting said audio data or said image data;

[1265] means for performing preprocessing such as noise removal and volume adjustment of audio data, or background removal and color correction of image data;

[1266] a means for analyzing the pre-processed data using a generative AI model; and

[1267] means for generating specific improvements based on the analysis results;

[1268] means for providing said improvements to a user;

[1269] A system including:

[1270] (Claim 2)

[1271] 10. The system of claim 1, further comprising means for pre-processing the audio data using FFT to perform noise removal and volume adjustment.

[1272] (Claim 3)

[1273] 10. The system of claim 1, further comprising means for pre-processing the image data using OpenCV for background removal and color correction.

[1274] "Application Example 1"

[1275] Claiming a new invention

[1276] (Claim 1)

[1277] means for receiving audio data or image data;

[1278] means for preprocessing the audio data or the image data;

[1279] A means of analyzing the pre-processed data using a generative AI model; and

[1280] means for generating specific improvements based on the analysis results;

[1281] means for providing said improvements to a user;

[1282] means for uploading data to a cloud server and receiving feedback;

[1283] means for installation on a smartphone or head-mounted display;

[1284] A system including:

[1285] (Claim 2)

[1286] 10. The system of claim 1, further comprising means for pre-processing the audio data to perform noise reduction and volume adjustment.

[1287] (Claim 3)

[1288] 10. The system of claim 1, further comprising means for performing background removal and color correction in pre-processing of the image data.

[1289] keyword:

[1290] Generative AI model, prompt sentence

[1291] "Example 2: Combining Emotion Engines"

[1292] (Claim 1)

[1293] means for receiving audio data or image data;

[1294] means for preprocessing the audio data or the image data;

[1295] A means of analyzing the pre-processed data using a generative AI model; and

[1296] means for generating specific improvements based on the analysis results and emotion recognition results by the emotion engine;

[1297] means for providing the user with advice according to the improvements and the user's emotional state;

[1298] means for modifying data based on said provided improvements and advice;

[1299] means for returning the corrected data and feedback information to the user;

[1300] A system including:

[1301] (Claim 2)

[1302] 10. The system of claim 1, further comprising means for pre-processing the audio data to perform noise reduction and volume adjustment.

[1303] (Claim 3)

[1304] 10. The system of claim 1, further comprising means for performing background removal and color correction in pre-processing of the image data.

[1305] "Application example 2 when combining emotion engines"

[1306] (Claim 1)

[1307] means for receiving audio data or image data;

[1308] means for preprocessing the audio data or the image data;

[1309] A means of analyzing the pre-processed data using a generative AI model; and

[1310] means for generating specific improvements based on the analysis results;

[1311] means for providing said improvements to a user;

[1312] A means for capturing work video and audio in real time using the camera and microphone of the smart glasses and uploading the data to a server;

[1313] A means for generating feedback according to the improvement points and emotions of the work in the server and displaying the feedback on a display of the smart glasses;

[1314] A system including:

[1315] (Claim 2)

[1316] 10. The system of claim 1, further comprising means for pre-processing the audio data to perform noise reduction and volume adjustment.

[1317] (Claim 3)

[1318] 10. The system of claim 1, further comprising means for performing background removal and color correction in pre-processing of the image data. [Explanation of symbols]

[1319] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving audio data or image data; means for preprocessing the audio data or the image data; a means of analyzing the pre-processed data using a generative AI model; and means for generating specific improvements based on the analysis results; means for providing said improvements to a user; A technology improvement support system including:

2. 2. The skill improvement support system according to claim 1, further comprising means for performing noise removal and volume adjustment in preprocessing of the audio data.

3. 2. The technology improvement support system according to claim 1, further comprising means for performing background removal and color correction in preprocessing of the image data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A