system

The system addresses the challenge of matching users with different sign language skills and interests by automating event planning and real-time translation, improving cross-cultural communication and emotional interaction.

JP2026070258APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing systems fail to efficiently match users with different sign language learning levels and interests, leading to restricted cross-cultural communication and educational opportunities, and lack sufficient real-time translation capabilities for smooth international communication.

Method used

A system that automatically matches users based on their sign language learning levels and interests, plans exchange events, and provides real-time translation between sign language and spoken language, using AI engines for translation and emotion recognition.

Benefits of technology

Facilitates smooth international communication by effectively matching users, planning events, and providing real-time translation, enhancing user experience and understanding through emotional feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070258000001_ABST
    Figure 2026070258000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] To facilitate international communication online, a means of automatically matching users with different sign language learning levels and interests is needed. A means to automatically plan and implement interaction events for matched users, A means of providing real-time, two-way translation between sign language and spoken language to support communication between users, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In international communication, there are language barriers among people with different sign languages, learning levels, and interests, making smooth communication difficult. As a result, there is a problem that cross-cultural communication and educational opportunities are restricted. In conventional systems, the efficient matching between users and the real-time translation function are insufficient, so the opportunities for cross-cultural communication cannot be provided sufficiently.

Means for Solving the Problems

[0005] This invention provides a means for automatically matching users with different sign language learning levels and interests to promote international sign language communication online. It also includes a means for automatically planning and conducting exchange events for the matched users. In addition, by using a means for real-time two-way translation between sign language and spoken language to support communication between users, it removes language barriers and enables smooth international exchange.

[0006] "Online" is a concept that refers to activities and connections conducted via the internet.

[0007] "International" refers to activities and exchanges that involve multiple countries and take place across national borders.

[0008] "Communication" refers to the process of sharing information and ideas with others.

[0009] "Sign language" generally refers to the visual, gesture-based means of communication used daily by people with hearing impairments.

[0010] "Learning level" refers to the stage an individual is currently at in acquiring a particular skill or knowledge.

[0011] "Interest" refers to the degree to which an individual has concern for or involvement with a particular matter or activity.

[0012] "Matching" refers to the process of combining two or more elements that are related or compatible based on specific criteria.

[0013] "Automatic" refers to a state in which a system operates according to set conditions and rules without requiring manual intervention.

[0014] An "exchange event" refers to an activity or opportunity where people exchange information and experiences.

[0015] "Planning" refers to the process of formulating plans or concepts to achieve specific goals.

[0016] "Implementation" refers to the act of actually putting planned activities or projects into action.

[0017] "Means" refers to the methods or processes used to achieve a certain goal.

[0018] "Translation" refers to the process of appropriately converting the content expressed in one language into another language.

[0019] "Bidirectional" refers to a state that enables information and influence to flow in two directions.

Brief Explanation of Drawings

[0020] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [[ID=?]] [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. <^ [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. <^ [Figure 9] It shows an emotion map to which multiple emotions are mapped. It seems there is a small formatting issue in the original text where line breaks might be a bit inconsistent in the sense of how they are used to separate paragraphs or sections. Also, there is an unclear tag <^ [Figure 8] and <^ [Figure 9] which might be a typo. I've translated it as best as possible while maintaining the integrity of the original text structure. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0021] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0022] First, the terms used in the following description will be described. [[ID=2�]]

[0023] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0024] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0025] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0026] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0028] [First Embodiment]

[0029] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0030] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0031] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0032] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0033] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0035] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0036] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0037] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0038] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0039] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0040] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0041] This invention is a system that facilitates online international sign language communication. The system is configured to effectively match individuals with different sign language learning levels and interests, and to automatically plan and conduct exchange events. It also provides real-time translation between sign language and spoken language, facilitating smooth communication.

[0042] The system works as follows: Users register by entering profile information such as their learning level and interests. The server then stores the user's information in a database. Based on this stored information, the server periodically scans users and matches them with other users who share similar interests and learning levels. For example, if a user sets their level to "intermediate" and "interested in movies," they can be matched with other users who meet the same criteria.

[0043] After matching, the server automatically schedules the networking event and generates a participation link. The device then notifies the user of these details and encourages them to participate. This allows for centralized scheduling and enables users to prepare smoothly.

[0044] During interaction, an AI engine is activated to support real-time translation of sign language and spoken language. This enables communication even if participants speak different languages. The device displays the translated results on the screen and outputs them as audio to convey information to other participants. For example, if a user says "I like sports" in sign language, it can be translated into spoken language and conveyed to other users as "I like sports."

[0045] Finally, after the event, the server provides users with a form to collect feedback. This feedback data will be used to improve the system in the future. Users can easily enter and send feedback from their own devices to the server. In this way, the system provides a comprehensive method for facilitating international sign language communication.

[0046] The following describes the processing flow.

[0047] Step 1:

[0048] Users access the system's website or application and open the registration form. They enter their name, email address, sign language proficiency level, and areas of interest, then click the register button.

[0049] Step 2:

[0050] The server receives registration information submitted by users and stores the information in a database. This is where each user's profile is established within the system.

[0051] Step 3:

[0052] The server periodically scans the database to identify users with matching learning levels and interests. It then uses an appropriate matching algorithm to form user groups with common attributes.

[0053] Step 4:

[0054] The server schedules online interaction events for matched users and generates participation links. This information is saved as an event schedule.

[0055] Step 5:

[0056] The server notifies the user's device of the generated event information. The device then displays the date of the social gathering and a participation link to the user. The user can then decide whether or not to participate.

[0057] Step 6:

[0058] When the exchange meeting begins, the device launches the video call app and establishes a video connection with other participants. While the video stream is established, users participate using sign language.

[0059] Step 7:

[0060] The server's AI engine analyzes the video stream and translates sign language into spoken language. The terminal displays this translation result on the screen in real time and outputs it as audio through its speaker.

[0061] Step 8:

[0062] After the exchange meeting ends, the server sends a feedback form to the participants. Users can evaluate the quality of the exchange by entering their opinions and impressions into this form and submitting it to the server.

[0063] Step 9:

[0064] The server stores the collected feedback in a database and uses it as material for future system improvements. This information contributes to improving the user experience.

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] To promote international dialogue, effectively matching users with diverse learning stages and interests and facilitating smooth exchange activities is a challenge. In particular, improving the accuracy of real-time translation of languages ​​and sign language is essential to facilitate communication among users. Furthermore, it is necessary to appropriately collect participant feedback obtained from these exchanges and utilize it for system improvement.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes means for automatically matching users with different learning stages and interests to facilitate online international dialogue; means for generating information for participation in exchange events and transmitting it to users' information terminals to inform them of the details of the dialogue; and means for performing real-time bidirectional translation between sign language and spoken language to support dialogue between users. This enables the smooth implementation of international exchange events and facilitates fluent dialogue even in diverse language environments.

[0070] "Online international dialogue" is a system that allows people living in different countries to communicate via the internet without meeting face-to-face.

[0071] "Learning stage" is an indicator that shows the level of proficiency and knowledge a user has in the process of learning sign language or other languages.

[0072] "Interest" refers to the interest or engagement that users have with a particular theme or activity.

[0073] "Matching" is the process of connecting users with common interests and learning stages, and identifying suitable partners for each individual.

[0074] "Interaction activities" refer to actions and events that deepen relationships and understanding among users through dialogue and collaborative work.

[0075] Sign language is a method of communication that does not involve sound, and it is a language that conveys meaning using the movements of the hands and fingers.

[0076] "Spoken language" refers to language transmitted through oral speech, and is a means of communication mediated by sound.

[0077] "Real-time translation" is a process that instantly converts input sign language or speech into another language, providing a translation service to the user without any time lag.

[0078] An "information terminal" is a device that can electronically process and display information and communicate, and generally includes smartphones and computers.

[0079] A "generative AI model" is a program or framework that uses artificial intelligence to analyze data and produce new information or results.

[0080] This invention is a system for facilitating international online communication. This system effectively matches users with diverse learning stages and interests, plans exchange events, and provides real-time translation between sign language and spoken language, thereby supporting communication among users.

[0081] The system is primarily composed of three components: a server, terminals, and users. The server stores users' registration information in a database and matches users with common interests and learning stages. Once a match is made, the server automatically schedules an exchange event and generates a participation link.

[0082] The device has a function to notify users of generated interaction event information. Furthermore, during interaction events, it utilizes a generation AI model to perform real-time translation between sign language and spoken language. The translated results are displayed on the device screen and are also output as audio.

[0083] Users can register their profile information via their device and participate in notified interaction events. After an event, they can also contribute to system improvement by easily entering and sending feedback to the server.

[0084] As a concrete example, if a user enters "intermediate sign language learner" and "likes traveling" in their profile, the server will use this information to match them with other users who have similar interests. If the user expresses "I'm looking forward to traveling" in sign language during an event, the generative AI model will translate it as "I look forward to traveling" and share it with other participants.

[0085] An example of a prompt is, "How will the user set up their profile information on the system to find the best sign language communication partner for self-learning?" Based on this prompt, the generating AI model identifies the optimal partner and performs effective matching.

[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0087] Step 1:

[0088] User Registration

[0089] Users input information about their sign language learning level and areas of interest through the system interface. This information is accurately stored in a database by the server. This creates a user profile that can be used in future matching processes. Specifically, users fill in the required data on the registration form and click the "Register" button.

[0090] Step 2:

[0091] Information preservation and organization

[0092] The server stores data received from users in a database and performs data cleaning to maintain consistency with existing data. Input data includes user profile information, and when stored in the database, this information is formatted and invalid data is removed. Specifically, SQL is used to add and update entries in the database.

[0093] Step 3:

[0094] Matching process

[0095] The server periodically scans the database and matches users based on common interests and learning levels. The input is the profiles of all users, which are used to execute the matching algorithm. The output generates information on matched user pairs. Specifically, this involves database searches using SQL queries and the application of the matching logic.

[0096] Step 4:

[0097] Event scheduling

[0098] The server selects an appropriate date based on the matching results and schedules the interaction event. The input is information about the matched users, and the output is the date of the configured event and a participation link. The specific operation includes the process of adjusting the date using the calendar API and generating the participation link.

[0099] Step 5:

[0100] Sending notifications

[0101] The terminal notifies the user of information about the generated interaction event. As input, it receives event information sent from the server and displays it on the user's screen via a notification mechanism. As output, detailed information for participating in the event is sent to the user. Specifically, this is done by sending a message via a push notification service.

[0102] Step 6:

[0103] Real-time translation

[0104] The device uses a generative AI model to translate sign language and spoken language during events. Inputs include sign language video and audio data captured by the camera, and output is translated text and audio. Specific operations include video analysis and speech synthesis processes.

[0105] Step 7:

[0106] Feedback Collection

[0107] Users can easily enter their thoughts and suggestions for improvement using a feedback form provided after the event. This input constitutes user feedback data, which is stored on the server and used for future improvements. Specifically, the feedback form is displayed, and a "submit" action is performed after data entry.

[0108] (Application Example 1)

[0109] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0110] Existing communication tools make it difficult for users of multiple nationalities to communicate information using sign language, and real-time information exchange between deaf and hearing-impaired users is also inconvenient. Therefore, there are challenges, particularly in situations like food delivery, where deaf individuals cannot communicate comfortably as users.

[0111] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0112] In this invention, the server includes means for automatically searching for users with different sign language learning levels and interests to facilitate multinational online information exchange, means for automatically planning and conducting meetings for matched users, and means for real-time conversion between sign language and spoken language to support information exchange between users. This enables deaf individuals to smoothly exchange information in real time, particularly facilitating mutual understanding between people in food delivery situations.

[0113] "Promoting multinational information exchange online" means enabling smooth information exchange among users with different linguistic backgrounds.

[0114] "Automatically searching for users with different sign language learning levels and interests" means that the system finds appropriate partners based on each user's sign language skills and interests.

[0115] "Automatically planning and executing meetings for matched users" means that the system automatically organizes and runs events where users with similar interests and skills can participate.

[0116] "Real-time conversion between sign language and spoken language to support information exchange between users" means translating what a user conveys in sign language into spoken language, and vice versa, thereby supporting communication between users of different languages.

[0117] "Using image input to convert sign language into speech and audio information into visual information" refers to analyzing sign language movements captured by a camera and converting them into speech, and conversely, replacing the speech with text or visual displays.

[0118] This invention provides a system that facilitates multinational information exchange and enables smooth communication, particularly between the hearing impaired and general users. This system is realized through the coordinated operation of three entities: a server, a terminal, and a user.

[0119] The server has a function to automatically search for users with different sign language learning levels and interests. Each user registers their sign language skills and interests as part of their profile information. Based on this information, the server identifies other users with similar profiles and performs matching.

[0120] Once matching is complete, the server automatically plans and executes the event. This process includes scheduling the event and generating participation links. This ensures that users can enjoy a consistently smooth event participation experience.

[0121] The user's device is equipped with hardware and software for real-time conversion between sign language and spoken language. Specifically, smartphones and tablets are used, and sign language is input as images using the camera and analyzed by an AI engine. The generative AI model used is GOOGLE TENSOR® FLOW®, and the Google® Cloud Speech-to-Text API is used for speech conversion. This makes it possible to convert from sign language to speech and from speech to visual information.

[0122] For example, if a user uses a food delivery service and tells the delivery person "add extra cheese" in sign language, this information will be converted in real time into a voice message saying "Please add extra cheese." Additionally, the delivery person's voice message "I've arrived" will be displayed on the user's device as either sign language or text.

[0123] Examples of prompt messages include the following:

[0124] "When a user says 'add extra cheese' in sign language, please output 'Please add extra cheese' verbally."

[0125] "Please create a feature that translates the delivery person's voice message 'Arrived' into sign language and displays it on the app."

[0126] In this way, the present invention provides a comprehensive method to support information transmission for the hearing impaired and improve their convenience.

[0127] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0128] Step 1:

[0129] Users input their sign language skills and interests into their terminal as part of their profile information. This information is sent directly to the server and stored in the database at high speed. The input here is the level of sign language ability and areas of interest, and the output is the user information stored on the server.

[0130] Step 2:

[0131] The server scans the database at regular intervals, automatically searching for and matching users with different sign language learning levels and interests. It uses an algorithm to identify other users with common interests or skills. The input for this step is the user information obtained in step 1, and the output is the matched user pairs.

[0132] Step 3:

[0133] Once a match is made, the server automatically plans the event and generates the schedule and participation link for the participants. These event details are sent to the device as notification data. Here, the event information based on the matching data is the input, and the event notification sent to the device is the output.

[0134] Step 4:

[0135] When converting between sign language and spoken language, the user's device uses its camera in real time to acquire sign language as image data, which is then analyzed using a generative AI model. The input is image data acquired by the camera, which is converted into sign language identification information and output.

[0136] Step 5:

[0137] The server performs real-time speech conversion based on the identification information and outputs it as audio data using the Google Cloud Speech-to-Text API. In this process, the sign language information identified from the image input is the input, and the audio data is the output.

[0138] Step 6:

[0139] Conversely, when audio is input to the device, the device analyzes it, converts it into text data, and then displays it on the screen as visual information. In this case, the audio data is the input, and the displayed text information or sign language-based visual information is the output.

[0140] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0141] This invention is a system for effectively facilitating online international sign language communication, primarily aimed at automatically matching users with different sign language learning levels and interests. The system also includes an emotion engine that analyzes users' facial expressions and voices to recognize emotions, further enhancing interaction between matched users.

[0142] The system operates as follows: Users register their profile information by accessing the system. This information is stored on the server and forms the basis for matching users with different learning levels and interests.

[0143] After a successful match is made, the server automatically plans an interaction event for the users and creates detailed event information. This information is sent from the server to each user's device, and users can prepare to participate in the event by receiving a notification.

[0144] During the interaction event, the device captures the user's facial expressions and voice via video call and sends them to the server's emotion engine. Based on the analyzed data, the server's emotion engine recognizes the user's emotions and generates appropriate feedback and actions.

[0145] For example, if a user smiles while using sign language, the emotion engine can detect the user's positive emotion and adjust the tone of communication accordingly. This allows other participants to understand that emotion and promotes better interaction.

[0146] In addition to real-time translation, the emotion engine further enhances the interaction experience. Emotion recognition results are recorded in a database and reflected in the user's profile. This information will be used when planning future events, enabling higher-quality matching.

[0147] After the event ends, the server distributes a feedback form to participants to collect opinions on the quality of interaction and areas for system improvement. Users input their feedback from their devices and send it to the server. This feedback is used for the continuous improvement of the system. As described above, the present invention improves the efficiency of communication and deepens understanding among participants.

[0148] The following describes the processing flow.

[0149] Step 1:

[0150] Users access the system's website or app and fill in the required fields on the registration form. This involves registering specific profile information, such as learning level and interests, with the understanding that this information will be shared with other users.

[0151] Step 2:

[0152] The server receives information about registered users and stores it in a database. This information is used in the matching process and forms the basis for supporting personalized experiences in the future.

[0153] Step 3:

[0154] The server periodically scans users based on stored data. It identifies and matches users with common interests and learning levels. Advanced algorithms are used here to achieve optimal pairing.

[0155] Step 4:

[0156] The server automatically generates event schedules for matched users and creates relevant participation information and links. This enables planned and efficient events.

[0157] Step 5:

[0158] The server sends the generated event information as a notification to each user's device. Upon receiving this notification, the device prompts the user to prepare for participation by presenting the event details.

[0159] Step 6:

[0160] When the exchange meeting begins, the device launches the video call application and establishes a video connection between participants. During this time, the device continuously acquires video, audio, and video data.

[0161] Step 7:

[0162] The device sends the acquired data to the server's emotion engine. The server analyzes this data, detecting facial expressions and voice tone to recognize the user's emotions.

[0163] Step 8:

[0164] The server's emotion engine adjusts the tone of communication and interface feedback as needed, based on real-time emotional information. This process is crucial for enriching the user experience.

[0165] Step 9:

[0166] After the interaction concludes, the server will distribute a feedback form to participants. Users can then submit their opinions about their experience during the event.

[0167] Step 10:

[0168] The server organizes and stores the collected feedback in a database, using it to improve future events. In this way, the system continuously improves and evolves into the optimal communication platform for users.

[0169] (Example 2)

[0170] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0171] As international communication is promoted, there is a need for smooth communication among users with different languages ​​and cultural backgrounds. However, appropriate pairing based on users' skill levels and interests is often not performed, and emotional communication is frequently lacking. A system is needed to solve these problems and realize deep understanding and smooth communication among users.

[0172] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0173] In this invention, the server includes means for automatically pairing users with different skill learning levels and interests to facilitate international online communication; means for automatically planning and conducting interaction activities for the paired users; and means for recognizing users' emotions using an emotion analysis function and generating feedback based on those emotions. This enables smooth communication and emotional transmission among diverse users.

[0174] "Means for facilitating international communication online" refers to a technological configuration that enables users with different linguistic and cultural backgrounds to interact effectively via the internet.

[0175] "Means for automatically pairing users with different skill learning levels and interests" refers to a technical function that selects and connects users with the most suitable partners based on their learning stage and interests, using their profile information.

[0176] "Means for automatically planning and implementing interaction activities" refers to a mechanism in which the system, after pairing users, schedules optimal interaction events and ensures that actual interactions take place.

[0177] "Feedback generation method using emotion analysis function" refers to a technical method for analyzing the user's facial expressions and voice data, understanding their emotions based on the results, and providing appropriate responses and adjustments.

[0178] "A means of providing real-time, two-way translation between voice and sign language to support communication between users" refers to a technology that instantly provides mutual conversion between voice and sign language to facilitate smooth communication between users using different communication methods.

[0179] This invention is a system that facilitates smooth interaction between users in order to promote international online communication. This system automatically pairs users based on different skill learning levels and interests, and supports communication using sentiment analysis.

[0180] The system's operation is primarily composed of three elements: server, terminal, and user. Users complete registration by entering their profile information from their terminal and sending it to the server. The server uses a generated AI model based on the collected data to optimally pair users and then plans the details of their interaction activities accordingly.

[0181] During video calls, the device collects the user's facial expressions and voice in real time and sends them to the server. The server uses an NVIDIA graphics card and EmotionAI software to analyze the facial expressions and voice and recognize the user's emotions. Based on the results of this emotion recognition, the server provides feedback and adjusts communication according to the user's emotions.

[0182] For example, if a user has an intermediate level of sign language and is interested in movies, the system uses this information to pair them with other users who have a similar skill level and interests. During interaction, if a user is conversing with an excited expression, the server's emotion engine detects that positive emotion and sends feedback to other participants so that they can also sense that emotion.

[0183] An example of a prompt might be: "Plan a movie-themed exchange event at an intermediate sign language level and suggest the best emotional feedback methods to promote positive emotions."

[0184] In this way, this system enables effective communication among diverse users and promotes deeper understanding and positive interaction among participants.

[0185] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0186] Step 1:

[0187] Users enter their profile information using a device. Specifically, they enter data such as their name, sign language proficiency level, and areas of interest. The entered information is sent from the device to the server. In this process, a data entry interface is used, and the entered information is converted into a digital form. The output is profile data used on the server.

[0188] Step 2:

[0189] The server stores the received profile data in the database. Security protocols are applied during data storage to ensure data consistency and security. The input is the user's profile data, which forms the basis for creating records in the database. The output is a secure database record.

[0190] Step 3:

[0191] The server utilizes a generative AI model to analyze profiles stored in a database. This analysis selects the optimal user pair based on different learning levels and interests. The input is the contents of the user database, and the output is a list of pairing candidates. The AI ​​model uses machine learning algorithms to learn from past successes and improve pairing accuracy.

[0192] Step 4:

[0193] The server plans appropriate interaction activities based on paired users. Specifically, it determines the date, time, and method of the interaction event and generates an event link based on each user's schedule. The input is pairing information, and the output is detailed information about the interaction event. The server executes a scheduling algorithm to select the optimal time and method.

[0194] Step 5:

[0195] The device uses video call functionality to collect the user's facial expressions and voice in real time during interaction activities. The collected data is sent to a server. The input is the user's video call data, and the output is the data sent to the server. The device uses a high-resolution camera and microphone to acquire the data.

[0196] Step 6:

[0197] The server analyzes the transmitted facial and audio data using EmotionAI software. Through this analysis, it recognizes the user's emotions and generates appropriate feedback. The input is video call data, and the output is the result of the emotion analysis and the feedback. Specific actions include executing a facial expression analysis algorithm.

[0198] Step 7:

[0199] Based on the analysis results, the server sends feedback tailored to the user's emotions to the terminal via the communication platform. The input is the result of the emotion analysis, and the output is the user's feedback message. The server uses a messaging API to deliver the feedback.

[0200] (Application Example 2)

[0201] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0202] In online international sign language communication, a challenge is to efficiently match users with different learning levels and interests, and to enrich their interactions. Furthermore, there is a need to provide real-time sign language sessions that take users' emotions into consideration, and to enable deeper learning and understanding through visual aids.

[0203] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0204] In this invention, the server includes means for automatically matching users with different sign language learning levels and interests to facilitate international online communication; means for automatically planning and conducting interaction events for the matched users; and means for analyzing the user's facial expressions and voice, recognizing their emotions, and adjusting the interaction experience based on the results. This makes it possible to provide interactive sign language sessions that take the user's emotions into account through visual devices such as smart glasses.

[0205] "Online international communication" refers to technology that enables the exchange of information across national borders using the internet.

[0206] "Sign language learning level and interests" refers to the level of proficiency in sign language knowledge and skills, as well as the specific areas of interest of each individual user.

[0207] "Automatic matching" refers to the process of mechanically pairing users together based on conditions that the system has set in advance.

[0208] "Automatically planning and implementing interaction events" refers to a process where a program automatically creates and executes a communication environment based on the combination of users.

[0209] "Real-time two-way translation" refers to a function that instantly translates different languages ​​during a conversation via an information processing device.

[0210] "Analyzing facial expressions and voice to recognize emotions" means using facial expression analysis software and voice recognition technology to analyze data on the user's facial expressions and voice to identify their emotions.

[0211] "Adjusting the interaction experience" refers to operations that optimize the flow of communication based on users' emotions and reactions, in order to provide a better experience.

[0212] "Receiving a sign language session through an information processing device" refers to receiving a sign language exchange conducted over a network using electronic devices.

[0213] "Displaying on a visual device" means displaying information on a vision-related device such as smart glasses or a display screen.

[0214] To realize this invention, three main elements—a server, a terminal, and a user—work together.

[0215] The server first maintains a database where users register their sign language learning level and interests. Once a user registers their profile, the server automatically matches them with other users based on this information and plans appropriate interaction events. The server also provides a real-time, two-way translation function between sign language and spoken language, facilitating smooth online communication.

[0216] The terminal primarily functions as a device for analyzing the user's emotions. Using visual devices such as smart glasses, the user's facial expressions and voice are analyzed using emotion recognition software (e.g., OpenFace), and the user's emotions are transmitted to a server. This emotion data is processed on the server, and the interaction experience is then refined.

[0217] Users can participate in sign language sessions and enjoy interactive communication with other users using smart glasses or information processing devices. Through smart glasses, they can visually receive others' sign language in real time and understand their emotions. For example, a scenario is envisioned where a sign language learner living in Japan has a real-time session with another learner in the United States to discuss environmental protection. In this case, the system can detect the user's passionate emotions, enabling a more interactive and deeper dialogue.

[0218] An example of a prompt could be: "Create an online environment where sign language learners from around the world can learn together. Design a system that recognizes each user's emotions and learning level and takes this into consideration during sign language sessions."

[0219] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0220] Step 1:

[0221] Users register their sign language learning level and interests using a terminal. The input data is the user's profile information, which is sent to the server. The server stores this data in a database, which forms the basis for matching.

[0222] Step 2:

[0223] The server automatically matches users based on stored profile information. The input data consists of multiple user profiles and their associated interests and learning levels. The server uses an algorithm to identify users with common interests and similar learning levels and generates an output that links their profiles.

[0224] Step 3:

[0225] The server automatically plans interaction events for matched users and sends notifications to their devices. The inputs are the information of the matched users and the matching results. The server automatically generates event details and provides output to the users' devices notifying them of the schedule and how to participate.

[0226] Step 4:

[0227] The device uses visual devices such as smart glasses to collect the user's facial expressions and voice. The input consists of real-time facial expression and voice data. Emotion recognition software is used to analyze this data and generate an output that identifies the user's emotions.

[0228] Step 5:

[0229] The server receives emotional data sent from the terminal and adjusts the interaction experience. The input is the result of the emotional analysis. Based on this result, the server adjusts the progress of the communication and generates output that provides appropriate feedback to the user.

[0230] Step 6:

[0231] Users receive real-time sign language sessions and interact with other users through smart glasses. Inputs include sign language video data from other users and adjustment feedback from the server. Outputs include the user's own learning improvement and enhanced communication, which the device displays visually.

[0232] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0233] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0234] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0235] [Second Embodiment]

[0236] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0237] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0238] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0239] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0240] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0241] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0242] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0243] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0244] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0245] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0246] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0247] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0248] This invention is a system that facilitates online international sign language communication. The system is configured to effectively match individuals with different sign language learning levels and interests, and to automatically plan and conduct exchange events. It also provides real-time translation between sign language and spoken language, facilitating smooth communication.

[0249] The system works as follows: Users register by entering profile information such as their learning level and interests. The server then stores the user's information in a database. Based on this stored information, the server periodically scans users and matches them with other users who share similar interests and learning levels. For example, if a user sets their level to "intermediate" and "interested in movies," they can be matched with other users who meet the same criteria.

[0250] After matching, the server automatically schedules the networking event and generates a participation link. The device then notifies the user of these details and encourages them to participate. This allows for centralized scheduling and enables users to prepare smoothly.

[0251] During interaction, an AI engine is activated to support real-time translation of sign language and spoken language. This enables communication even if participants speak different languages. The device displays the translated results on the screen and outputs them as audio to convey information to other participants. For example, if a user says "I like sports" in sign language, it can be translated into spoken language and conveyed to other users as "I like sports."

[0252] Finally, after the event, the server provides users with a form to collect feedback. This feedback data will be used to improve the system in the future. Users can easily enter and send feedback from their own devices to the server. In this way, the system provides a comprehensive method for facilitating international sign language communication.

[0253] The following describes the processing flow.

[0254] Step 1:

[0255] Users access the system's website or application and open the registration form. They enter their name, email address, sign language proficiency level, and areas of interest, then click the register button.

[0256] Step 2:

[0257] The server receives registration information submitted by users and stores the information in a database. This is where each user's profile is established within the system.

[0258] Step 3:

[0259] The server periodically scans the database to identify users with matching learning levels and interests. It then uses an appropriate matching algorithm to form user groups with common attributes.

[0260] Step 4:

[0261] The server schedules online interaction events for matched users and generates participation links. This information is saved as an event schedule.

[0262] Step 5:

[0263] The server notifies the user's device of the generated event information. The device then displays the date of the social gathering and a participation link to the user. The user can then decide whether or not to participate.

[0264] Step 6:

[0265] When the exchange meeting begins, the device launches the video call app and establishes a video connection with other participants. While the video stream is established, users participate using sign language.

[0266] Step 7:

[0267] The server's AI engine analyzes the video stream and translates sign language into spoken language. The terminal displays this translation result on the screen in real time and outputs it as audio through its speaker.

[0268] Step 8:

[0269] After the exchange meeting ends, the server sends a feedback form to the participants. Users can evaluate the quality of the exchange by entering their opinions and impressions into this form and submitting it to the server.

[0270] Step 9:

[0271] The server stores the collected feedback in a database and uses it as material for future system improvements. This information contributes to improving the user experience.

[0272] (Example 1)

[0273] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0274] To promote international dialogue, effectively matching users with diverse learning stages and interests and facilitating smooth exchange activities is a challenge. In particular, improving the accuracy of real-time translation of languages ​​and sign language is essential to facilitate communication among users. Furthermore, it is necessary to appropriately collect participant feedback obtained from these exchanges and utilize it for system improvement.

[0275] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0276] In this invention, the server includes means for automatically matching users with different learning stages and interests to promote online international conversations, means for generating information for participating in communication events and transmitting it to the users' information terminals to notify the details of the conversations, and means for performing real-time two-way translation between sign language and spoken language to assist the conversations between users. As a result, international communication events can be smoothly carried out, and fluent conversations are possible even in diverse language environments.

[0277] "Online international conversation" refers to a mechanism through which people living in different countries communicate with each other without direct face-to-face contact via the Internet.

[0278] "Learning stage" refers to an indicator that shows the degree of proficiency and knowledge in the process of users learning sign language or language.

[0279] "Interest" means the interest and enthusiasm that users have for specific themes or activities.

[0280] "Matching" is a process of combining users with common interests or learning stages and identifying suitable partners for each.

[0281] "Communication activity" refers to actions and events for deepening relationships and understanding through conversations and joint work carried out among users.

[0282] "Sign language" is a communication method without voice, a language that conveys meaning using the movements of hands and fingers.

[0283] "Spoken language" is a language transmitted by oral vocalization, a communication means via voice.

[0284] "Real-time translation" is a process of immediately converting the input sign language or voice into another language, a translation service provided to users without a time lag.

[0285] An "information terminal" is a device that can electronically process and display information and communicate, and generally includes smartphones and computers.

[0286] A "generative AI model" is a program or framework that utilizes artificial intelligence to analyze data and produce new information and results.

[0287] The present invention is a system for smoothly conducting international online conversations. This system effectively matches users with various learning levels and interests, plans communication events, and provides real-time translation between sign language and spoken language to support conversations between users.

[0288] The system is mainly composed of three entities: a server, a terminal, and a user. The server stores the registration information of users in a database and matches users with common interests and learning levels. When a match is established, the server automatically sets the schedule of the communication event and generates a participation link.

[0289] The terminal has a function of notifying users of the generated communication event information. Also, during the communication event, it utilizes the generative AI model to perform real-time translation between sign language and spoken language. The translated result is displayed on the screen of the terminal and voice output is also performed.

[0290] Users can register their profile information via the terminal and participate in the notified communication events. Also, after the event ends, they can contribute to the improvement of the system by simply inputting feedback and sending it to the server.

[0291] As a specific example, when a user enters "intermediate sign language user" and "likes traveling" in their profile, the server matches other users with similar interests based on this information. If the user expresses "I'm looking forward to traveling" in sign language during the event, it will be translated to "I look forward to traveling" through the generative AI model and conveyed to other participants.

[0292] An example of a prompt is, "How will the user set up their profile information on the system to find the best sign language communication partner for self-learning?" Based on this prompt, the generating AI model identifies the optimal partner and performs effective matching.

[0293] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0294] Step 1:

[0295] User Registration

[0296] Users input information about their sign language learning level and areas of interest through the system interface. This information is accurately stored in a database by the server. This creates a user profile that can be used in future matching processes. Specifically, users fill in the required data on the registration form and click the "Register" button.

[0297] Step 2:

[0298] Information preservation and organization

[0299] The server stores data received from users in a database and performs data cleaning to maintain consistency with existing data. Input data includes user profile information, and when stored in the database, this information is formatted and invalid data is removed. Specifically, SQL is used to add and update entries in the database.

[0300] Step 3:

[0301] Matching process

[0302] The server scans the database regularly and matches users based on common interests and learning levels. The input is the profiles of all users, and based on this, the matching algorithm is executed. As output, pair information of the matched users is generated. As specific operations, database search by SQL query and application of matching logic are performed.

[0303] Step 4:

[0304] Event Scheduling

[0305] The server selects an appropriate schedule based on the matching results and schedules communication events. The input is the information of the matched users, and the output is the schedule of the set events and participation links. Specific operations include the process of adjusting the schedule using the calendar API and generating participation links.

[0306] Step 5:

[0307] Notification Sending

[0308] The terminal notifies the user of the information of the generated communication event. As input, it receives the event information sent from the server and displays this on the user's screen through the notification mechanism. As output, detailed information for event participation is sent to the user. As specific operations, message sending via the PUSH notification service is performed.

[0309] Step 6:

[0310] Real - Time Translation

[0311] The terminal utilizes the generated AI model during the event to perform sign language and spoken language translation. As input, sign language video and audio data captured by the camera are used, and as output, translated text and audio are generated. Specific operations include the process of video analysis and speech synthesis.

[0312] Step 7:

[0313] Feedback Collection

[0314] Users can easily enter their thoughts and suggestions for improvement using a feedback form provided after the event. This input constitutes user feedback data, which is stored on the server and used for future improvements. Specifically, the feedback form is displayed, and a "submit" action is performed after data entry.

[0315] (Application Example 1)

[0316] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0317] Existing communication tools make it difficult for users of multiple nationalities to communicate information using sign language, and real-time information exchange between deaf and hearing-impaired users is also inconvenient. Therefore, there are challenges, particularly in situations like food delivery, where deaf individuals cannot communicate comfortably as users.

[0318] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0319] In this invention, the server includes means for automatically searching for users with different sign language learning levels and interests to facilitate multinational online information exchange, means for automatically planning and conducting meetings for matched users, and means for real-time conversion between sign language and spoken language to support information exchange between users. This enables deaf individuals to smoothly exchange information in real time, particularly facilitating mutual understanding between people in food delivery situations.

[0320] "Promoting multinational information exchange online" means enabling smooth information exchange among users with different linguistic backgrounds.

[0321] "Automatically searching for users with different sign language learning levels and interests" means that the system finds appropriate partners based on each user's sign language skills and interests.

[0322] "Automatically planning and executing meetings for matched users" means that the system automatically organizes and runs events where users with similar interests and skills can participate.

[0323] "Real-time conversion between sign language and spoken language to support information exchange between users" means translating what a user conveys in sign language into spoken language, and vice versa, thereby supporting communication between users of different languages.

[0324] "Using image input to convert sign language into speech and audio information into visual information" refers to analyzing sign language movements captured by a camera and converting them into speech, and conversely, replacing the speech with text or visual displays.

[0325] This invention provides a system that facilitates multinational information exchange and enables smooth communication, particularly between the hearing impaired and general users. This system is realized through the coordinated operation of three entities: a server, a terminal, and a user.

[0326] The server has a function to automatically search for users with different sign language learning levels and interests. Each user registers their sign language skills and interests as part of their profile information. Based on this information, the server identifies other users with similar profiles and performs matching.

[0327] Once matching is complete, the server automatically plans and executes the event. This process includes scheduling the event and generating participation links. This ensures that users can enjoy a consistently smooth event participation experience.

[0328] The user's device is equipped with hardware and software for real-time conversion between sign language and spoken language. Specifically, smartphones and tablets are used, with the camera used to input sign language as images, which are then analyzed by an AI engine. Google TensorFlow is used as the generative AI model, and the Google Cloud Speech-to-Text API is used for speech conversion. This enables conversion from sign language to speech and from speech to visual information.

[0329] For example, if a user uses a food delivery service and tells the delivery person "add extra cheese" in sign language, this information will be converted in real time into a voice message saying "Please add extra cheese." Additionally, the delivery person's voice message "I've arrived" will be displayed on the user's device as either sign language or text.

[0330] Examples of prompt messages include the following:

[0331] "When a user says 'add extra cheese' in sign language, please output 'Please add extra cheese' verbally."

[0332] "Please create a feature that translates the delivery person's voice message 'Arrived' into sign language and displays it on the app."

[0333] In this way, the present invention provides a comprehensive method to support information transmission for the hearing impaired and improve their convenience.

[0334] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0335] Step 1:

[0336] Users input their sign language skills and interests into their terminal as part of their profile information. This information is sent directly to the server and stored in the database at high speed. The input here is the level of sign language ability and areas of interest, and the output is the user information stored on the server.

[0337] Step 2:

[0338] The server scans the database at regular intervals, automatically searching for and matching users with different sign language learning levels and interests. It uses an algorithm to identify other users with common interests or skills. The input for this step is the user information obtained in step 1, and the output is the matched user pairs.

[0339] Step 3:

[0340] Once a match is made, the server automatically plans the event and generates the schedule and participation link for the participants. These event details are sent to the device as notification data. Here, the event information based on the matching data is the input, and the event notification sent to the device is the output.

[0341] Step 4:

[0342] When converting between sign language and spoken language, the user's device uses its camera in real time to acquire sign language as image data, which is then analyzed using a generative AI model. The input is image data acquired by the camera, which is converted into sign language identification information and output.

[0343] Step 5:

[0344] The server performs real-time speech conversion based on the identification information and outputs it as audio data using the Google Cloud Speech-to-Text API. In this process, the sign language information identified from the image input is the input, and the audio data is the output.

[0345] Step 6:

[0346] Conversely, when audio is input to the device, the device analyzes it, converts it into text data, and then displays it on the screen as visual information. In this case, the audio data is the input, and the displayed text information or sign language-based visual information is the output.

[0347] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0348] This invention is a system for effectively facilitating online international sign language communication, primarily aimed at automatically matching users with different sign language learning levels and interests. The system also includes an emotion engine that analyzes users' facial expressions and voices to recognize emotions, further enhancing interaction between matched users.

[0349] The system operates as follows: Users register their profile information by accessing the system. This information is stored on the server and forms the basis for matching users with different learning levels and interests.

[0350] After a successful match is made, the server automatically plans an interaction event for the users and creates detailed event information. This information is sent from the server to each user's device, and users can prepare to participate in the event by receiving a notification.

[0351] During the interaction event, the device captures the user's facial expressions and voice via video call and sends them to the server's emotion engine. Based on the analyzed data, the server's emotion engine recognizes the user's emotions and generates appropriate feedback and actions.

[0352] For example, if a user smiles while using sign language, the emotion engine can detect the user's positive emotion and adjust the tone of communication accordingly. This allows other participants to understand that emotion and promotes better interaction.

[0353] In addition to real-time translation, the emotion engine further enhances the interaction experience. Emotion recognition results are recorded in a database and reflected in the user's profile. This information will be used when planning future events, enabling higher-quality matching.

[0354] After the event ends, the server distributes a feedback form to participants to collect opinions on the quality of interaction and areas for system improvement. Users input their feedback from their devices and send it to the server. This feedback is used for the continuous improvement of the system. As described above, the present invention improves the efficiency of communication and deepens understanding among participants.

[0355] The following describes the processing flow.

[0356] Step 1:

[0357] Users access the system's website or app and fill in the required fields on the registration form. This involves registering specific profile information, such as learning level and interests, with the understanding that this information will be shared with other users.

[0358] Step 2:

[0359] The server receives information about registered users and stores it in a database. This information is used in the matching process and forms the basis for supporting personalized experiences in the future.

[0360] Step 3:

[0361] The server periodically scans users based on stored data. It identifies and matches users with common interests and learning levels. Advanced algorithms are used here to achieve optimal pairing.

[0362] Step 4:

[0363] The server automatically generates event schedules for matched users and creates relevant participation information and links. This enables planned and efficient events.

[0364] Step 5:

[0365] The server sends the generated event information as a notification to each user's device. Upon receiving this notification, the device prompts the user to prepare for participation by presenting the event details.

[0366] Step 6:

[0367] When the exchange meeting begins, the device launches the video call application and establishes a video connection between participants. During this time, the device continuously acquires video, audio, and video data.

[0368] Step 7:

[0369] The device sends the acquired data to the server's emotion engine. The server analyzes this data, detecting facial expressions and voice tone to recognize the user's emotions.

[0370] Step 8:

[0371] The server's emotion engine adjusts the tone of communication and interface feedback as needed, based on real-time emotional information. This process is crucial for enriching the user experience.

[0372] Step 9:

[0373] After the interaction concludes, the server will distribute a feedback form to participants. Users can then submit their opinions about their experience during the event.

[0374] Step 10:

[0375] The server organizes and stores the collected feedback in a database, using it to improve future events. In this way, the system continuously improves and evolves into the optimal communication platform for users.

[0376] (Example 2)

[0377] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0378] As international communication is promoted, there is a need for smooth communication among users with different languages ​​and cultural backgrounds. However, appropriate pairing based on users' skill levels and interests is often not performed, and emotional communication is frequently lacking. A system is needed to solve these problems and realize deep understanding and smooth communication among users.

[0379] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0380] In this invention, the server includes means for automatically pairing users with different skill learning levels and interests to facilitate international online communication; means for automatically planning and conducting interaction activities for the paired users; and means for recognizing users' emotions using an emotion analysis function and generating feedback based on those emotions. This enables smooth communication and emotional transmission among diverse users.

[0381] "Means for facilitating international communication online" refers to a technological configuration that enables users with different linguistic and cultural backgrounds to interact effectively via the internet.

[0382] "Means for automatically pairing users with different skill learning levels and interests" refers to a technical function that selects and connects users with the most suitable partners based on their learning stage and interests, using their profile information.

[0383] "Means for automatically planning and implementing interaction activities" refers to a mechanism in which the system, after pairing users, schedules optimal interaction events and ensures that actual interactions take place.

[0384] "Feedback generation method using emotion analysis function" refers to a technical method for analyzing the user's facial expressions and voice data, understanding their emotions based on the results, and providing appropriate responses and adjustments.

[0385] "A means of providing real-time, two-way translation between voice and sign language to support communication between users" refers to a technology that instantly provides mutual conversion between voice and sign language to facilitate smooth communication between users using different communication methods.

[0386] This invention is a system that facilitates smooth interaction between users in order to promote international online communication. This system automatically pairs users based on different skill learning levels and interests, and supports communication using sentiment analysis.

[0387] The system's operation is primarily composed of three elements: server, terminal, and user. Users complete registration by entering their profile information from their terminal and sending it to the server. The server uses a generated AI model based on the collected data to optimally pair users and then plans the details of their interaction activities accordingly.

[0388] During video calls, the device collects the user's facial expressions and voice in real time and sends them to the server. The server uses an NVIDIA graphics card and EmotionAI software to analyze the facial expressions and voice and recognize the user's emotions. Based on the results of this emotion recognition, the server provides feedback and adjusts communication according to the user's emotions.

[0389] For example, if a user has an intermediate level of sign language and is interested in movies, the system uses this information to pair them with other users who have a similar skill level and interests. During interaction, if a user is conversing with an excited expression, the server's emotion engine detects that positive emotion and sends feedback to other participants so that they can also sense that emotion.

[0390] An example of a prompt might be: "Plan a movie-themed exchange event at an intermediate sign language level and suggest the best emotional feedback methods to promote positive emotions."

[0391] In this way, this system enables effective communication among diverse users and promotes deeper understanding and positive interaction among participants.

[0392] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0393] Step 1:

[0394] Users enter their profile information using a device. Specifically, they enter data such as their name, sign language proficiency level, and areas of interest. The entered information is sent from the device to the server. In this process, a data entry interface is used, and the entered information is converted into a digital form. The output is profile data used on the server.

[0395] Step 2:

[0396] The server stores the received profile data in the database. Security protocols are applied during data storage to ensure data consistency and security. The input is the user's profile data, which forms the basis for creating records in the database. The output is a secure database record.

[0397] Step 3:

[0398] The server utilizes a generative AI model to analyze profiles stored in a database. This analysis selects the optimal user pair based on different learning levels and interests. The input is the contents of the user database, and the output is a list of pairing candidates. The AI ​​model uses machine learning algorithms to learn from past successes and improve pairing accuracy.

[0399] Step 4:

[0400] The server plans appropriate interaction activities based on paired users. Specifically, it determines the date, time, and method of the interaction event and generates an event link based on each user's schedule. The input is pairing information, and the output is detailed information about the interaction event. The server executes a scheduling algorithm to select the optimal time and method.

[0401] Step 5:

[0402] The device uses video call functionality to collect the user's facial expressions and voice in real time during interaction activities. The collected data is sent to a server. The input is the user's video call data, and the output is the data sent to the server. The device uses a high-resolution camera and microphone to acquire the data.

[0403] Step 6:

[0404] The server analyzes the transmitted facial and audio data using EmotionAI software. Through this analysis, it recognizes the user's emotions and generates appropriate feedback. The input is video call data, and the output is the result of the emotion analysis and the feedback. Specific actions include executing a facial expression analysis algorithm.

[0405] Step 7:

[0406] Based on the analysis results, the server sends feedback tailored to the user's emotions to the terminal via the communication platform. The input is the result of the emotion analysis, and the output is the user's feedback message. The server uses a messaging API to deliver the feedback.

[0407] (Application Example 2)

[0408] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0409] In online international sign language communication, a challenge is to efficiently match users with different learning levels and interests, and to enrich their interactions. Furthermore, there is a need to provide real-time sign language sessions that take users' emotions into consideration, and to enable deeper learning and understanding through visual aids.

[0410] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0411] In this invention, the server includes means for automatically matching users with different sign language learning levels and interests to facilitate international online communication; means for automatically planning and conducting interaction events for the matched users; and means for analyzing the user's facial expressions and voice, recognizing their emotions, and adjusting the interaction experience based on the results. This makes it possible to provide interactive sign language sessions that take the user's emotions into account through visual devices such as smart glasses.

[0412] "Online international communication" refers to technology that enables the exchange of information across national borders using the internet.

[0413] "Sign language learning level and interests" refers to the level of proficiency in sign language knowledge and skills, as well as the specific areas of interest of each individual user.

[0414] "Automatic matching" refers to the process of mechanically pairing users together based on conditions that the system has set in advance.

[0415] "Automatically planning and implementing interaction events" refers to a process where a program automatically creates and executes a communication environment based on the combination of users.

[0416] "Real-time two-way translation" refers to a function that instantly translates different languages ​​during a conversation via an information processing device.

[0417] "Analyzing facial expressions and voice to recognize emotions" means using facial expression analysis software and voice recognition technology to analyze data on the user's facial expressions and voice to identify their emotions.

[0418] "Adjusting the interaction experience" refers to operations that optimize the flow of communication based on users' emotions and reactions, in order to provide a better experience.

[0419] "Receiving a sign language session through an information processing device" refers to receiving a sign language exchange conducted over a network using electronic devices.

[0420] "Displaying on a visual device" means displaying information on a vision-related device such as smart glasses or a display screen.

[0421] To realize this invention, three main elements—a server, a terminal, and a user—work together.

[0422] The server first maintains a database where users register their sign language learning level and interests. Once a user registers their profile, the server automatically matches them with other users based on this information and plans appropriate interaction events. The server also provides a real-time, two-way translation function between sign language and spoken language, facilitating smooth online communication.

[0423] The terminal primarily functions as a device for analyzing the user's emotions. Using visual devices such as smart glasses, the user's facial expressions and voice are analyzed using emotion recognition software (e.g., OpenFace), and the user's emotions are transmitted to a server. This emotion data is processed on the server, and the interaction experience is then refined.

[0424] Users can participate in sign language sessions and enjoy interactive communication with other users using smart glasses or information processing devices. Through smart glasses, they can visually receive others' sign language in real time and understand their emotions. For example, a scenario is envisioned where a sign language learner living in Japan has a real-time session with another learner in the United States to discuss environmental protection. In this case, the system can detect the user's passionate emotions, enabling a more interactive and deeper dialogue.

[0425] An example of a prompt could be: "Create an online environment where sign language learners from around the world can learn together. Design a system that recognizes each user's emotions and learning level and takes this into consideration during sign language sessions."

[0426] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0427] Step 1:

[0428] Users register their sign language learning level and interests using a terminal. The input data is the user's profile information, which is sent to the server. The server stores this data in a database, which forms the basis for matching.

[0429] Step 2:

[0430] The server automatically matches users based on stored profile information. The input data consists of multiple user profiles and their associated interests and learning levels. The server uses an algorithm to identify users with common interests and similar learning levels and generates an output that links their profiles.

[0431] Step 3:

[0432] The server automatically plans interaction events for matched users and sends notifications to their devices. The inputs are the information of the matched users and the matching results. The server automatically generates event details and provides output to the users' devices notifying them of the schedule and how to participate.

[0433] Step 4:

[0434] The device uses visual devices such as smart glasses to collect the user's facial expressions and voice. The input consists of real-time facial expression and voice data. Emotion recognition software is used to analyze this data and generate an output that identifies the user's emotions.

[0435] Step 5:

[0436] The server receives emotional data sent from the terminal and adjusts the interaction experience. The input is the result of the emotional analysis. Based on this result, the server adjusts the progress of the communication and generates output that provides appropriate feedback to the user.

[0437] Step 6:

[0438] Users receive real-time sign language sessions and interact with other users through smart glasses. Inputs include sign language video data from other users and adjustment feedback from the server. Outputs include the user's own learning improvement and enhanced communication, which the device displays visually.

[0439] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0440] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0441] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0442] [Third Embodiment]

[0443] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0444] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0445] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0446] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0447] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0448] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0449] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0450] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0451] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0452] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0453] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0454] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0455] This invention is a system that facilitates online international sign language communication. The system is configured to effectively match individuals with different sign language learning levels and interests, and to automatically plan and conduct exchange events. It also provides real-time translation between sign language and spoken language, facilitating smooth communication.

[0456] The system works as follows: Users register by entering profile information such as their learning level and interests. The server then stores the user's information in a database. Based on this stored information, the server periodically scans users and matches them with other users who share similar interests and learning levels. For example, if a user sets their level to "intermediate" and "interested in movies," they can be matched with other users who meet the same criteria.

[0457] After matching, the server automatically schedules the networking event and generates a participation link. The device then notifies the user of these details and encourages them to participate. This allows for centralized scheduling and enables users to prepare smoothly.

[0458] During interaction, an AI engine is activated to support real-time translation of sign language and spoken language. This enables communication even if participants speak different languages. The device displays the translated results on the screen and outputs them as audio to convey information to other participants. For example, if a user says "I like sports" in sign language, it can be translated into spoken language and conveyed to other users as "I like sports."

[0459] Finally, after the event, the server provides users with a form to collect feedback. This feedback data will be used to improve the system in the future. Users can easily enter and send feedback from their own devices to the server. In this way, the system provides a comprehensive method for facilitating international sign language communication.

[0460] The following describes the processing flow.

[0461] Step 1:

[0462] Users access the system's website or application and open the registration form. They enter their name, email address, sign language proficiency level, and areas of interest, then click the register button.

[0463] Step 2:

[0464] The server receives registration information submitted by users and stores the information in a database. This is where each user's profile is established within the system.

[0465] Step 3:

[0466] The server periodically scans the database to identify users with matching learning levels and interests. It then uses an appropriate matching algorithm to form user groups with common attributes.

[0467] Step 4:

[0468] The server schedules online interaction events for matched users and generates participation links. This information is saved as an event schedule.

[0469] Step 5:

[0470] The server notifies the user's device of the generated event information. The device then displays the date of the social gathering and a participation link to the user. The user can then decide whether or not to participate.

[0471] Step 6:

[0472] When the exchange meeting begins, the device launches the video call app and establishes a video connection with other participants. While the video stream is established, users participate using sign language.

[0473] Step 7:

[0474] The server's AI engine analyzes the video stream and translates sign language into spoken language. The terminal displays this translation result on the screen in real time and outputs it as audio through its speaker.

[0475] Step 8:

[0476] After the exchange meeting ends, the server sends a feedback form to the participants. Users can evaluate the quality of the exchange by entering their opinions and impressions into this form and submitting it to the server.

[0477] Step 9:

[0478] The server stores the collected feedback in a database and uses it as material for future system improvements. This information contributes to improving the user experience.

[0479] (Example 1)

[0480] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0481] To promote international dialogue, effectively matching users with diverse learning stages and interests and facilitating smooth exchange activities is a challenge. In particular, improving the accuracy of real-time translation of languages ​​and sign language is essential to facilitate communication among users. Furthermore, it is necessary to appropriately collect participant feedback obtained from these exchanges and utilize it for system improvement.

[0482] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0483] In this invention, the server includes means for automatically matching users with different learning stages and interests to facilitate online international dialogue; means for generating information for participation in exchange events and transmitting it to users' information terminals to inform them of the details of the dialogue; and means for performing real-time bidirectional translation between sign language and spoken language to support dialogue between users. This enables the smooth implementation of international exchange events and facilitates fluent dialogue even in diverse language environments.

[0484] "Online international dialogue" is a system that allows people living in different countries to communicate via the internet without meeting face-to-face.

[0485] "Learning stage" is an indicator that shows the level of proficiency and knowledge a user has in the process of learning sign language or other languages.

[0486] "Interest" refers to the interest or engagement that users have with a particular theme or activity.

[0487] "Matching" is the process of connecting users with common interests and learning stages, and identifying suitable partners for each individual.

[0488] "Interaction activities" refer to actions and events that deepen relationships and understanding among users through dialogue and collaborative work.

[0489] Sign language is a method of communication that does not involve sound, and it is a language that conveys meaning using the movements of the hands and fingers.

[0490] "Spoken language" refers to language transmitted through oral speech, and is a means of communication mediated by sound.

[0491] "Real-time translation" is a process that instantly converts input sign language or speech into another language, providing a translation service to the user without any time lag.

[0492] An "information terminal" is a device that can electronically process and display information and communicate, and generally includes smartphones and computers.

[0493] A "generative AI model" is a program or framework that uses artificial intelligence to analyze data and produce new information or results.

[0494] This invention is a system for facilitating international online communication. This system effectively matches users with diverse learning stages and interests, plans exchange events, and provides real-time translation between sign language and spoken language, thereby supporting communication among users.

[0495] The system is primarily composed of three components: a server, terminals, and users. The server stores users' registration information in a database and matches users with common interests and learning stages. Once a match is made, the server automatically schedules an exchange event and generates a participation link.

[0496] The device has a function to notify users of generated interaction event information. Furthermore, during interaction events, it utilizes a generation AI model to perform real-time translation between sign language and spoken language. The translated results are displayed on the device screen and also output as audio.

[0497] Users can register their profile information via their device and participate in notified interaction events. After an event, they can also contribute to system improvement by easily entering and sending feedback to the server.

[0498] As a concrete example, if a user enters "intermediate sign language learner" and "likes traveling" in their profile, the server will use this information to match them with other users who have similar interests. If the user expresses "I'm looking forward to traveling" in sign language during an event, the generative AI model will translate it as "I look forward to traveling" and share it with other participants.

[0499] An example of a prompt is, "How will the user set up their profile information on the system to find the best sign language communication partner for self-learning?" Based on this prompt, the generating AI model identifies the most suitable partner and performs effective matching.

[0500] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0501] Step 1:

[0502] User Registration

[0503] Users input information about their sign language learning level and areas of interest through the system interface. This information is accurately stored in a database by the server. This creates a user profile that can be used in future matching processes. Specifically, users fill in the required data on the registration form and click the "Register" button.

[0504] Step 2:

[0505] Information preservation and organization

[0506] The server stores data received from users in a database and performs data cleaning to maintain consistency with existing data. Input data includes user profile information, and when stored in the database, this information is formatted and invalid data is removed. Specifically, SQL is used to add and update entries in the database.

[0507] Step 3:

[0508] Matching process

[0509] The server periodically scans the database and matches users based on common interests and learning levels. The input is the profiles of all users, which are used to execute the matching algorithm. The output generates information on matched user pairs. Specifically, this involves database searches using SQL queries and the application of the matching logic.

[0510] Step 4:

[0511] Event scheduling

[0512] The server selects an appropriate date based on the matching results and schedules the interaction event. The input is information about the matched users, and the output is the date of the configured event and a participation link. The specific operation includes the process of adjusting the date using the calendar API and generating the participation link.

[0513] Step 5:

[0514] Sending notifications

[0515] The terminal notifies the user of information about the generated interaction event. As input, it receives event information sent from the server and displays it on the user's screen via a notification mechanism. As output, detailed information for participating in the event is sent to the user. Specifically, this is done by sending a message via a push notification service.

[0516] Step 6:

[0517] Real-time translation

[0518] The device uses a generative AI model to translate sign language and spoken language during events. Inputs include sign language video and audio data captured by the camera, and output is translated text and audio. Specific operations include video analysis and speech synthesis processes.

[0519] Step 7:

[0520] Feedback Collection

[0521] Users can easily enter their thoughts and suggestions for improvement using a feedback form provided after the event ends. This input constitutes user feedback data, which is stored on the server and used for future improvements. Specifically, the feedback form is displayed, and a "submit" action is performed after data entry.

[0522] (Application Example 1)

[0523] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0524] Existing communication tools make it difficult for users of multiple nationalities to communicate information using sign language, and real-time information exchange between deaf and hearing-impaired users is also inconvenient. Therefore, there are challenges, particularly in situations like food delivery, where deaf individuals cannot communicate comfortably as users.

[0525] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0526] In this invention, the server includes means for automatically searching for users with different sign language learning levels and interests to facilitate multinational online information exchange, means for automatically planning and conducting meetings for matched users, and means for real-time conversion between sign language and spoken language to support information exchange between users. This enables deaf individuals to smoothly exchange information in real time, particularly facilitating mutual understanding between people in food delivery situations.

[0527] "Promoting multinational information exchange online" means enabling smooth information exchange among users with different linguistic backgrounds.

[0528] "Automatically searching for users with different sign language learning levels and interests" means that the system finds appropriate partners based on each user's sign language skills and interests.

[0529] "Automatically planning and executing meetings for matched users" means that the system automatically organizes and runs events where users with similar interests and skills can participate.

[0530] "Real-time conversion between sign language and spoken language to support information exchange between users" means translating what a user conveys in sign language into spoken language, and vice versa, thereby supporting communication between users of different languages.

[0531] "Using image input to convert sign language into speech and audio information into visual information" refers to analyzing sign language movements captured by a camera and converting them into speech, and conversely, replacing the speech with text or visual displays.

[0532] This invention provides a system that facilitates multinational information exchange and enables smooth communication, particularly between the hearing impaired and general users. This system is realized through the coordinated operation of three entities: a server, a terminal, and a user.

[0533] The server has a function to automatically search for users with different sign language learning levels and interests. Each user registers their sign language skills and interests as part of their profile information. Based on this information, the server identifies other users with similar profiles and performs matching.

[0534] Once matching is complete, the server automatically plans and executes the event. This process includes scheduling the event and generating participation links. This ensures that users can enjoy a consistently smooth event participation experience.

[0535] The user's device is equipped with hardware and software for real-time conversion between sign language and spoken language. Specifically, smartphones and tablets are used, with the camera used to input sign language as images, which are then analyzed by an AI engine. Google TensorFlow is used as the generative AI model, and the Google Cloud Speech-to-Text API is used for speech conversion. This enables conversion from sign language to speech and from speech to visual information.

[0536] For example, if a user uses a food delivery service and tells the delivery person "add extra cheese" in sign language, this information will be converted in real time into a voice message saying "Please add extra cheese." Additionally, the delivery person's voice message "I've arrived" will be displayed on the user's device as either sign language or text.

[0537] Examples of prompt messages include the following:

[0538] "When a user says 'add extra cheese' in sign language, please output 'Please add extra cheese' verbally."

[0539] "Please create a feature that translates the delivery person's voice message 'Arrived' into sign language and displays it on the app."

[0540] In this way, the present invention provides a comprehensive method to support information transmission for the hearing impaired and improve their convenience.

[0541] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0542] Step 1:

[0543] Users input their sign language skills and interests into their terminal as part of their profile information. This information is sent directly to the server and stored in the database at high speed. The input here is the level of sign language ability and areas of interest, and the output is the user information stored on the server.

[0544] Step 2:

[0545] The server scans the database at regular intervals, automatically searching for and matching users with different sign language learning levels and interests. It uses an algorithm to identify other users with common interests or skills. The input for this step is the user information obtained in step 1, and the output is the matched user pairs.

[0546] Step 3:

[0547] Once a match is made, the server automatically plans the event and generates the schedule and participation link for the participants. These event details are sent to the device as notification data. Here, the event information based on the matching data is the input, and the event notification sent to the device is the output.

[0548] Step 4:

[0549] When converting between sign language and spoken language, the user's device uses its camera in real time to acquire sign language as image data, which is then analyzed using a generative AI model. The input is image data acquired by the camera, which is converted into sign language identification information and output.

[0550] Step 5:

[0551] The server performs real-time speech conversion based on the identification information and outputs it as audio data using the Google Cloud Speech-to-Text API. In this process, the sign language information identified from the image input is the input, and the audio data is the output.

[0552] Step 6:

[0553] Conversely, when audio is input to the device, the device analyzes it, converts it into text data, and then displays it on the screen as visual information. In this case, the audio data is the input, and the displayed text information or sign language-based visual information is the output.

[0554] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0555] This invention is a system for effectively facilitating online international sign language communication, primarily aimed at automatically matching users with different sign language learning levels and interests. The system also includes an emotion engine that analyzes users' facial expressions and voices to recognize emotions, further enhancing interaction between matched users.

[0556] The system operates as follows: Users register their profile information by accessing the system. This information is stored on the server and forms the basis for matching users with different learning levels and interests.

[0557] After a successful match is made, the server automatically plans an interaction event for the users and creates detailed event information. This information is sent from the server to each user's device, and users can prepare to participate in the event by receiving a notification.

[0558] During the interaction event, the device captures the user's facial expressions and voice via video call and sends them to the server's emotion engine. Based on the analyzed data, the server's emotion engine recognizes the user's emotions and generates appropriate feedback and actions.

[0559] For example, if a user smiles while using sign language, the emotion engine can detect the user's positive emotion and adjust the tone of communication accordingly. This allows other participants to understand that emotion and promotes better interaction.

[0560] In addition to real-time translation, the emotion engine further enhances the interaction experience. Emotion recognition results are recorded in a database and reflected in the user's profile. This information will be used when planning future events, enabling higher-quality matching.

[0561] After the event ends, the server distributes a feedback form to participants to collect opinions on the quality of interaction and areas for system improvement. Users input their feedback from their devices and send it to the server. This feedback is used for the continuous improvement of the system. As described above, the present invention improves the efficiency of communication and deepens understanding among participants.

[0562] The following describes the processing flow.

[0563] Step 1:

[0564] Users access the system's website or app and fill in the required fields on the registration form. This involves registering specific profile information, such as learning level and interests, with the understanding that this information will be shared with other users.

[0565] Step 2:

[0566] The server receives information about registered users and stores it in a database. This information is used in the matching process and forms the basis for supporting personalized experiences in the future.

[0567] Step 3:

[0568] The server periodically scans users based on stored data. It identifies and matches users with common interests and learning levels. Advanced algorithms are used here to achieve optimal pairing.

[0569] Step 4:

[0570] The server automatically generates event schedules for matched users and creates relevant participation information and links. This enables planned and efficient events.

[0571] Step 5:

[0572] The server sends the generated event information as a notification to each user's device. Upon receiving this notification, the device prompts the user to prepare for participation by presenting the event details.

[0573] Step 6:

[0574] When the exchange meeting begins, the device launches the video call application and establishes a video connection between participants. During this time, the device continuously acquires video, audio, and video data.

[0575] Step 7:

[0576] The device sends the acquired data to the server's emotion engine. The server analyzes this data, detecting facial expressions and voice tone to recognize the user's emotions.

[0577] Step 8:

[0578] The server's emotion engine adjusts the tone of communication and interface feedback as needed, based on real-time emotional information. This process is crucial for enriching the user experience.

[0579] Step 9:

[0580] After the interaction concludes, the server will distribute a feedback form to participants. Users can then submit their opinions about their experience during the event.

[0581] Step 10:

[0582] The server organizes and stores the collected feedback in a database, using it to improve future events. In this way, the system continuously improves and evolves into the optimal communication platform for users.

[0583] (Example 2)

[0584] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0585] As international communication is promoted, there is a need for smooth communication among users with different languages ​​and cultural backgrounds. However, appropriate pairing based on users' skill levels and interests is often not performed, and emotional communication is frequently lacking. A system is needed to solve these problems and realize deep understanding and smooth communication among users.

[0586] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0587] In this invention, the server includes means for automatically pairing users with different skill learning levels and interests to facilitate international online communication; means for automatically planning and conducting interaction activities for the paired users; and means for recognizing users' emotions using an emotion analysis function and generating feedback based on those emotions. This enables smooth communication and emotional transmission among diverse users.

[0588] "Means for facilitating international communication online" refers to a technological configuration that enables users with different linguistic and cultural backgrounds to interact effectively via the internet.

[0589] "Means for automatically pairing users with different skill learning levels and interests" refers to a technical function that selects and connects users with the most suitable partners based on their learning stage and interests, using their profile information.

[0590] "Means for automatically planning and implementing interaction activities" refers to a mechanism in which the system, after pairing users, schedules optimal interaction events and ensures that actual interactions take place.

[0591] "Feedback generation method using emotion analysis function" refers to a technical method for analyzing the user's facial expressions and voice data, understanding their emotions based on the results, and providing appropriate responses and adjustments.

[0592] "A means of providing real-time, two-way translation between voice and sign language to support communication between users" refers to a technology that instantly provides mutual conversion between voice and sign language to facilitate smooth communication between users using different communication methods.

[0593] This invention is a system that facilitates smooth interaction between users in order to promote international online communication. This system automatically pairs users based on different skill learning levels and interests, and supports communication using sentiment analysis.

[0594] The system's operation is primarily composed of three elements: server, terminal, and user. Users complete registration by entering their profile information from their terminal and sending it to the server. The server uses a generated AI model based on the collected data to optimally pair users and then plans the details of their interaction activities accordingly.

[0595] During video calls, the device collects the user's facial expressions and voice in real time and sends them to the server. The server uses an NVIDIA graphics card and EmotionAI software to analyze the facial expressions and voice and recognize the user's emotions. Based on the results of this emotion recognition, the server provides feedback and adjusts communication according to the user's emotions.

[0596] For example, if a user has an intermediate level of sign language and is interested in movies, the system uses this information to pair them with other users who have a similar skill level and interests. During interaction, if a user is conversing with an excited expression, the server's emotion engine detects that positive emotion and sends feedback to other participants so that they can also sense that emotion.

[0597] An example of a prompt might be: "Plan a movie-themed exchange event at an intermediate sign language level and suggest the best emotional feedback methods to promote positive emotions."

[0598] In this way, this system enables effective communication among diverse users and promotes deeper understanding and positive interaction among participants.

[0599] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0600] Step 1:

[0601] Users enter their profile information using a device. Specifically, they enter data such as their name, sign language proficiency level, and areas of interest. The entered information is sent from the device to the server. In this process, a data entry interface is used, and the entered information is converted into a digital form. The output is profile data used on the server.

[0602] Step 2:

[0603] The server stores the received profile data in the database. Security protocols are applied during data storage to ensure data consistency and security. The input is the user's profile data, which forms the basis for creating records in the database. The output is a secure database record.

[0604] Step 3:

[0605] The server utilizes a generative AI model to analyze profiles stored in a database. This analysis selects the optimal user pair based on different learning levels and interests. The input is the contents of the user database, and the output is a list of pairing candidates. The AI ​​model uses machine learning algorithms to learn from past successes and improve pairing accuracy.

[0606] Step 4:

[0607] The server plans appropriate interaction activities based on paired users. Specifically, it determines the date, time, and method of the interaction event and generates an event link based on each user's schedule. The input is pairing information, and the output is detailed information about the interaction event. The server executes a scheduling algorithm to select the optimal time and method.

[0608] Step 5:

[0609] The device uses video call functionality to collect the user's facial expressions and voice in real time during interaction activities. The collected data is sent to a server. The input is the user's video call data, and the output is the data sent to the server. The device uses a high-resolution camera and microphone to acquire the data.

[0610] Step 6:

[0611] The server analyzes the transmitted facial and audio data using EmotionAI software. Through this analysis, it recognizes the user's emotions and generates appropriate feedback. The input is video call data, and the output is the result of the emotion analysis and the feedback. Specific actions include executing a facial expression analysis algorithm.

[0612] Step 7:

[0613] Based on the analysis results, the server sends feedback tailored to the user's emotions to the terminal via the communication platform. The input is the result of the emotion analysis, and the output is the user's feedback message. The server uses a messaging API to deliver the feedback.

[0614] (Application Example 2)

[0615] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0616] In online international sign language communication, a challenge is to efficiently match users with different learning levels and interests, and to enrich their interactions. Furthermore, there is a need to provide real-time sign language sessions that take users' emotions into consideration, and to enable deeper learning and understanding through visual aids.

[0617] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0618] In this invention, the server includes means for automatically matching users with different sign language learning levels and interests to facilitate international online communication; means for automatically planning and conducting interaction events for the matched users; and means for analyzing the user's facial expressions and voice, recognizing their emotions, and adjusting the interaction experience based on the results. This makes it possible to provide interactive sign language sessions that take the user's emotions into account through visual devices such as smart glasses.

[0619] "Online international communication" refers to technology that enables the exchange of information across national borders using the internet.

[0620] "Sign language learning level and interests" refers to the level of proficiency in sign language knowledge and skills, as well as the specific areas of interest of each individual user.

[0621] "Automatic matching" refers to a process where the system mechanically pairs users together based on pre-set conditions.

[0622] "Automatically planning and implementing interaction events" means that the program automatically creates and executes a communication space based on the combination of users.

[0623] "Real-time two-way translation" refers to a function that instantly translates different languages ​​during a conversation via an information processing device.

[0624] "Analyzing facial expressions and voice to recognize emotions" means using facial expression analysis software and voice recognition technology to analyze data on the user's facial expressions and voice to identify their emotions.

[0625] "Adjusting the interaction experience" refers to operations that optimize the flow of communication based on users' emotions and reactions, in order to provide a better experience.

[0626] "Receiving a sign language session through an information processing device" refers to receiving a sign language exchange conducted over a network using electronic devices.

[0627] "Displaying on a visual device" means displaying information on a vision-related device such as smart glasses or a display screen.

[0628] To realize this invention, three main elements—a server, a terminal, and a user—work together.

[0629] The server first maintains a database where users register their sign language learning level and interests. Once a user registers their profile, the server automatically matches them with other users based on this information and plans appropriate interaction events. The server also provides a real-time, two-way translation function between sign language and spoken language, facilitating smooth online communication.

[0630] The terminal primarily functions as a device for analyzing the user's emotions. Using visual devices such as smart glasses, the user's facial expressions and voice are analyzed using emotion recognition software (e.g., OpenFace), and the user's emotions are transmitted to a server. This emotion data is processed on the server, and the interaction experience is then refined.

[0631] Users can participate in sign language sessions and enjoy interactive communication with other users using smart glasses or information processing devices. Through smart glasses, they can visually receive others' sign language in real time and understand their emotions. For example, a scenario is envisioned where a sign language learner living in Japan has a real-time session with another learner in the United States to discuss environmental protection. In this case, the system can detect the user's passionate emotions, enabling a more interactive and deeper dialogue.

[0632] As an example of a prompt, you could use the request: "Create an online environment where sign language learners from around the world can learn together. Design a system that recognizes each user's emotions and learning level and takes this into consideration during sign language sessions."

[0633] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0634] Step 1:

[0635] Users register their sign language learning level and interests using a terminal. The input data is the user's profile information, which is sent to the server. The server stores this data in a database, which forms the basis for matching.

[0636] Step 2:

[0637] The server automatically matches users based on stored profile information. The input data consists of multiple user profiles and their associated interests and learning levels. The server uses an algorithm to identify users with common interests and similar learning levels and generates an output that links their profiles.

[0638] Step 3:

[0639] The server automatically plans interaction events for matched users and sends notifications to their devices. The inputs are the information of the matched users and the matching results. The server automatically generates event details and provides output to the users' devices notifying them of the schedule and how to participate.

[0640] Step 4:

[0641] The device uses visual devices such as smart glasses to collect the user's facial expressions and voice. The input consists of real-time facial expression and voice data. Emotion recognition software is used to analyze this data and generate an output that identifies the user's emotions.

[0642] Step 5:

[0643] The server receives emotional data sent from the terminal and adjusts the interaction experience. The input is the result of the emotional analysis. Based on this result, the server adjusts the progress of the communication and generates output that provides appropriate feedback to the user.

[0644] Step 6:

[0645] Users receive real-time sign language sessions and interact with other users through smart glasses. Inputs include sign language video data from other users and adjustment feedback from the server. Outputs include the user's own learning improvement and enhanced communication, which the device displays visually.

[0646] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0647] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0648] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0649] [Fourth Embodiment]

[0650] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0651] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0652] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0653] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0654] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0655] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0656] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0657] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0658] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0659] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0660] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0661] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0662] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0663] This invention is a system that facilitates online international sign language communication. The system is configured to effectively match individuals with different sign language learning levels and interests, and to automatically plan and conduct exchange events. It also provides real-time translation between sign language and spoken language, facilitating smooth communication.

[0664] The system works as follows: Users register by entering profile information such as their learning level and interests. The server then stores the user's information in a database. Based on this stored information, the server periodically scans users and matches them with other users who share similar interests and learning levels. For example, if a user sets their level to "intermediate" and "interested in movies," they can be matched with other users who meet the same criteria.

[0665] After matching, the server automatically schedules the networking event and generates a participation link. The device then notifies the user of these details and encourages them to participate. This allows for centralized scheduling and enables users to prepare smoothly.

[0666] During interaction, an AI engine is activated to support real-time translation of sign language and spoken language. This enables communication even if participants speak different languages. The device displays the translated results on the screen and outputs them as audio to convey information to other participants. For example, if a user says "I like sports" in sign language, it can be translated into spoken language and conveyed to other users as "I like sports."

[0667] Finally, after the event, the server provides users with a form to collect feedback. This feedback data will be used to improve the system in the future. Users can easily enter and send feedback from their own devices to the server. In this way, the system provides a comprehensive method for facilitating international sign language communication.

[0668] The following describes the processing flow.

[0669] Step 1:

[0670] Users access the system's website or application and open the registration form. They enter their name, email address, sign language proficiency level, and areas of interest, then click the register button.

[0671] Step 2:

[0672] The server receives registration information submitted by users and stores the information in a database. This is where each user's profile is established within the system.

[0673] Step 3:

[0674] The server periodically scans the database to identify users with matching learning levels and interests. It then uses an appropriate matching algorithm to form user groups with common attributes.

[0675] Step 4:

[0676] The server schedules online interaction events for matched users and generates participation links. This information is saved as an event schedule.

[0677] Step 5:

[0678] The server notifies the user's device of the generated event information. The device then displays the date of the social gathering and a participation link to the user. The user can then decide whether or not to participate.

[0679] Step 6:

[0680] When the exchange meeting begins, the device launches the video call app and establishes a video connection with other participants. While the video stream is established, users participate using sign language.

[0681] Step 7:

[0682] The server's AI engine analyzes the video stream and translates sign language into spoken language. The terminal displays this translation result on the screen in real time and outputs it as audio through its speaker.

[0683] Step 8:

[0684] After the exchange meeting ends, the server sends a feedback form to the participants. Users can evaluate the quality of the exchange by entering their opinions and impressions into this form and submitting it to the server.

[0685] Step 9:

[0686] The server stores the collected feedback in a database and uses it as material for future system improvements. This information contributes to improving the user experience.

[0687] (Example 1)

[0688] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0689] To promote international dialogue, effectively matching users with diverse learning stages and interests and facilitating smooth exchange activities is a challenge. In particular, improving the accuracy of real-time translation of languages ​​and sign language is essential to facilitate communication among users. Furthermore, it is necessary to appropriately collect participant feedback obtained from these exchanges and utilize it for system improvement.

[0690] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0691] In this invention, the server includes means for automatically matching users with different learning stages and interests to facilitate online international dialogue; means for generating information for participation in exchange events and transmitting it to users' information terminals to inform them of the details of the dialogue; and means for performing real-time bidirectional translation between sign language and spoken language to support dialogue between users. This enables the smooth implementation of international exchange events and facilitates fluent dialogue even in diverse language environments.

[0692] "Online international dialogue" is a system that allows people living in different countries to communicate via the internet without meeting face-to-face.

[0693] "Learning stage" is an indicator that shows the level of proficiency and knowledge a user has in the process of learning sign language or other languages.

[0694] "Interest" refers to the interest or engagement that users have with a particular theme or activity.

[0695] "Matching" is the process of connecting users with common interests and learning stages, and identifying suitable partners for each individual.

[0696] "Interaction activities" refer to actions and events that deepen relationships and understanding among users through dialogue and collaborative work.

[0697] Sign language is a method of communication that does not involve sound, and it is a language that conveys meaning using the movements of the hands and fingers.

[0698] "Spoken language" refers to language transmitted through oral speech, and is a means of communication mediated by sound.

[0699] "Real-time translation" is a process that instantly converts input sign language or speech into another language, providing a translation service to the user without any time lag.

[0700] An "information terminal" is a device that can electronically process and display information and communicate, and generally includes smartphones and computers.

[0701] A "generative AI model" is a program or framework that uses artificial intelligence to analyze data and produce new information or results.

[0702] This invention is a system for facilitating international online communication. This system effectively matches users with diverse learning stages and interests, plans exchange events, and provides real-time translation between sign language and spoken language, thereby supporting communication among users.

[0703] The system is primarily composed of three components: a server, terminals, and users. The server stores users' registration information in a database and matches users with common interests and learning stages. Once a match is made, the server automatically schedules an exchange event and generates a participation link.

[0704] The device has a function to notify users of generated interaction event information. Furthermore, during interaction events, it utilizes a generation AI model to perform real-time translation between sign language and spoken language. The translated results are displayed on the device screen and also output as audio.

[0705] Users can register their profile information via their device and participate in notified interaction events. After an event, they can also contribute to system improvement by easily entering and sending feedback to the server.

[0706] As a concrete example, if a user enters "intermediate sign language learner" and "likes traveling" in their profile, the server will use this information to match them with other users who have similar interests. If the user expresses "I'm looking forward to traveling" in sign language during an event, the generative AI model will translate it as "I look forward to traveling" and share it with other participants.

[0707] An example of a prompt is, "How will the user set up their profile information on the system to find the best sign language communication partner for self-learning?" Based on this prompt, the generating AI model identifies the most suitable partner and performs effective matching.

[0708] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0709] Step 1:

[0710] User Registration

[0711] Users input information about their sign language learning level and areas of interest through the system interface. This information is accurately stored in a database by the server. This creates a user profile that can be used in future matching processes. Specifically, users fill in the required data on the registration form and click the "Register" button.

[0712] Step 2:

[0713] Information preservation and organization

[0714] The server stores data received from users in a database and performs data cleaning to maintain consistency with existing data. Input data includes user profile information, and when stored in the database, this information is formatted and invalid data is removed. Specifically, SQL is used to add and update entries in the database.

[0715] Step 3:

[0716] Matching process

[0717] The server periodically scans the database and matches users based on common interests and learning levels. The input is the profiles of all users, which are used to execute the matching algorithm. The output generates information on matched user pairs. Specifically, this involves database searches using SQL queries and the application of the matching logic.

[0718] Step 4:

[0719] Event scheduling

[0720] The server selects an appropriate date based on the matching results and schedules the interaction event. The input is information about the matched users, and the output is the date of the configured event and a participation link. The specific operation includes the process of adjusting the date using the calendar API and generating the participation link.

[0721] Step 5:

[0722] Sending notifications

[0723] The terminal notifies the user of information about the generated interaction event. As input, it receives event information sent from the server and displays it on the user's screen via a notification mechanism. As output, detailed information for participating in the event is sent to the user. Specifically, this is done by sending a message via a push notification service.

[0724] Step 6:

[0725] Real-time translation

[0726] The device uses a generative AI model to translate sign language and spoken language during events. Inputs include sign language video and audio data captured by the camera, and output is translated text and audio. Specific operations include video analysis and speech synthesis processes.

[0727] Step 7:

[0728] Feedback Collection

[0729] Users can easily enter their thoughts and suggestions for improvement using a feedback form provided after the event ends. This input constitutes user feedback data, which is stored on the server and used for future improvements. Specifically, the feedback form is displayed, and a "submit" action is performed after data entry.

[0730] (Application Example 1)

[0731] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0732] Existing communication tools make it difficult for users of multiple nationalities to communicate information using sign language, and real-time information exchange between deaf and hearing-impaired users is also inconvenient. Therefore, there are challenges, particularly in situations like food delivery, where deaf individuals cannot communicate comfortably as users.

[0733] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0734] In this invention, the server includes means for automatically searching for users with different sign language learning levels and interests to facilitate multinational online information exchange, means for automatically planning and conducting meetings for matched users, and means for real-time conversion between sign language and spoken language to support information exchange between users. This enables deaf individuals to smoothly exchange information in real time, particularly facilitating mutual understanding between people in food delivery situations.

[0735] "Promoting multinational information exchange online" means enabling smooth information exchange among users with different linguistic backgrounds.

[0736] "Automatically searching for users with different sign language learning levels and interests" means that the system finds appropriate partners based on each user's sign language skills and interests.

[0737] "Automatically planning and executing meetings for matched users" means that the system automatically organizes and runs events where users with similar interests and skills can participate.

[0738] "Real-time conversion between sign language and spoken language to support information exchange between users" means translating what a user conveys in sign language into spoken language, and vice versa, thereby supporting communication between users of different languages.

[0739] "Using image input to convert sign language into speech and audio information into visual information" refers to analyzing sign language movements captured by a camera and converting them into speech, and conversely, replacing the speech with text or visual displays.

[0740] This invention provides a system that facilitates multinational information exchange and enables smooth communication, particularly between the hearing impaired and general users. This system is realized through the coordinated operation of three entities: a server, a terminal, and a user.

[0741] The server has a function to automatically search for users with different sign language learning levels and interests. Each user registers their sign language skills and interests as part of their profile information. Based on this information, the server identifies other users with similar profiles and performs matching.

[0742] Once matching is complete, the server automatically plans and executes the event. This process includes scheduling the event and generating participation links. This ensures that users can enjoy a consistently smooth event participation experience.

[0743] The user's device is equipped with hardware and software for real-time conversion between sign language and spoken language. Specifically, smartphones and tablets are used, with the camera used to input sign language as images, which are then analyzed by an AI engine. Google TensorFlow is used as the generative AI model, and the Google Cloud Speech-to-Text API is used for speech conversion. This enables conversion from sign language to speech and from speech to visual information.

[0744] For example, if a user uses a food delivery service and tells the delivery person "add extra cheese" in sign language, this information will be converted in real time into a voice message saying "Please add extra cheese." Additionally, the delivery person's voice message "I've arrived" will be displayed on the user's device as either sign language or text.

[0745] Examples of prompt messages include the following:

[0746] "When a user says 'add extra cheese' in sign language, please output 'Please add extra cheese' verbally."

[0747] "Please create a feature that translates the delivery person's voice message 'Arrived' into sign language and displays it on the app."

[0748] In this way, the present invention provides a comprehensive method to support information transmission for the hearing impaired and improve their convenience.

[0749] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0750] Step 1:

[0751] Users input their sign language skills and interests into their terminal as part of their profile information. This information is sent directly to the server and stored in the database at high speed. The input here is the level of sign language ability and areas of interest, and the output is the user information stored on the server.

[0752] Step 2:

[0753] The server scans the database at regular intervals, automatically searching for and matching users with different sign language learning levels and interests. It uses an algorithm to identify other users with common interests or skills. The input for this step is the user information obtained in step 1, and the output is the matched user pairs.

[0754] Step 3:

[0755] Once a match is made, the server automatically plans the event and generates the schedule and participation link for the participants. These event details are sent to the device as notification data. Here, the event information based on the matching data is the input, and the event notification sent to the device is the output.

[0756] Step 4:

[0757] When converting between sign language and spoken language, the user's device uses its camera in real time to acquire sign language as image data, which is then analyzed using a generative AI model. The input is image data acquired by the camera, which is converted into sign language identification information and output.

[0758] Step 5:

[0759] The server performs real-time speech conversion based on the identification information and outputs it as audio data using the Google Cloud Speech-to-Text API. In this process, the sign language information identified from the image input is the input, and the audio data is the output.

[0760] Step 6:

[0761] Conversely, when audio is input to the device, the device analyzes it, converts it into text data, and then displays it on the screen as visual information. In this case, the audio data is the input, and the displayed text information or sign language-based visual information is the output.

[0762] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0763] This invention is a system for effectively facilitating online international sign language communication, primarily aimed at automatically matching users with different sign language learning levels and interests. The system also includes an emotion engine that analyzes users' facial expressions and voices to recognize emotions, further enhancing interaction between matched users.

[0764] The system operates as follows: Users register their profile information by accessing the system. This information is stored on the server and forms the basis for matching users with different learning levels and interests.

[0765] After a successful match is made, the server automatically plans an interaction event for the users and creates detailed event information. This information is sent from the server to each user's device, and users can prepare to participate in the event by receiving a notification.

[0766] During the interaction event, the device captures the user's facial expressions and voice via video call and sends them to the server's emotion engine. Based on the analyzed data, the server's emotion engine recognizes the user's emotions and generates appropriate feedback and actions.

[0767] For example, if a user smiles while using sign language, the emotion engine can detect the user's positive emotion and adjust the tone of communication accordingly. This allows other participants to understand that emotion and promotes better interaction.

[0768] In addition to real-time translation, the emotion engine further enhances the interaction experience. Emotion recognition results are recorded in a database and reflected in the user's profile. This information will be used when planning future events, enabling higher-quality matching.

[0769] After the event ends, the server distributes a feedback form to participants to collect opinions on the quality of interaction and areas for system improvement. Users input their feedback from their devices and send it to the server. This feedback is used for the continuous improvement of the system. As described above, the present invention improves the efficiency of communication and deepens understanding among participants.

[0770] The following describes the processing flow.

[0771] Step 1:

[0772] Users access the system's website or app and fill in the required fields on the registration form. This involves registering specific profile information, such as learning level and interests, with the understanding that this information will be shared with other users.

[0773] Step 2:

[0774] The server receives information about registered users and stores it in a database. This information is used in the matching process and forms the basis for supporting personalized experiences in the future.

[0775] Step 3:

[0776] The server periodically scans users based on stored data. It identifies and matches users with common interests and learning levels. Advanced algorithms are used here to achieve optimal pairing.

[0777] Step 4:

[0778] The server automatically generates event schedules for matched users and creates relevant participation information and links. This enables planned and efficient events.

[0779] Step 5:

[0780] The server sends the generated event information as a notification to each user's device. Upon receiving this notification, the device prompts the user to prepare for participation by presenting the event details.

[0781] Step 6:

[0782] When the exchange meeting begins, the device launches the video call application and establishes a video connection between participants. During this time, the device continuously acquires video, audio, and video data.

[0783] Step 7:

[0784] The device sends the acquired data to the server's emotion engine. The server analyzes this data, detecting facial expressions and voice tone to recognize the user's emotions.

[0785] Step 8:

[0786] The server's emotion engine adjusts the tone of communication and interface feedback as needed, based on real-time emotional information. This process is crucial for enriching the user experience.

[0787] Step 9:

[0788] After the interaction concludes, the server will distribute a feedback form to participants. Users can then submit their opinions about their experience during the event.

[0789] Step 10:

[0790] The server organizes and stores the collected feedback in a database, using it to improve future events. In this way, the system continuously improves and evolves into the optimal communication platform for users.

[0791] (Example 2)

[0792] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0793] As international communication is promoted, there is a need for smooth communication among users with different languages ​​and cultural backgrounds. However, appropriate pairing based on users' skill levels and interests is often not performed, and emotional communication is frequently lacking. A system is needed to solve these problems and realize deep understanding and smooth communication among users.

[0794] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0795] In this invention, the server includes means for automatically pairing users with different skill learning levels and interests to facilitate international online communication; means for automatically planning and conducting interaction activities for the paired users; and means for recognizing users' emotions using an emotion analysis function and generating feedback based on those emotions. This enables smooth communication and emotional transmission among diverse users.

[0796] "Means for facilitating international communication online" refers to a technological configuration that enables users with different linguistic and cultural backgrounds to interact effectively via the internet.

[0797] "Means for automatically pairing users with different skill learning levels and interests" refers to a technical function that selects and connects users with the most suitable partners based on their learning stage and interests, using their profile information.

[0798] "Means for automatically planning and implementing interaction activities" refers to a mechanism in which the system, after pairing users, schedules optimal interaction events and ensures that actual interactions take place.

[0799] "Feedback generation method using emotion analysis function" refers to a technical method for analyzing the user's facial expressions and voice data, understanding their emotions based on the results, and providing appropriate responses and adjustments.

[0800] "A means of providing real-time, two-way translation between voice and sign language to support communication between users" refers to a technology that instantly provides mutual conversion between voice and sign language to facilitate smooth communication between users using different communication methods.

[0801] This invention is a system that facilitates smooth interaction between users in order to promote international online communication. This system automatically pairs users based on different skill learning levels and interests, and supports communication using sentiment analysis.

[0802] The system's operation is primarily composed of three elements: server, terminal, and user. Users complete registration by entering their profile information from their terminal and sending it to the server. The server uses a generated AI model based on the collected data to optimally pair users and then plans the details of their interaction activities accordingly.

[0803] During video calls, the device collects the user's facial expressions and voice in real time and sends them to the server. The server uses an NVIDIA graphics card and EmotionAI software to analyze the facial expressions and voice and recognize the user's emotions. Based on the results of this emotion recognition, the server provides feedback and adjusts communication according to the user's emotions.

[0804] For example, if a user has an intermediate level of sign language and is interested in movies, the system uses this information to pair them with other users who have a similar skill level and interests. During interaction, if a user is conversing with an excited expression, the server's emotion engine detects that positive emotion and sends feedback to other participants so that they can also sense that emotion.

[0805] An example of a prompt might be: "Plan a movie-themed exchange event at an intermediate sign language level and suggest the best emotional feedback methods to promote positive emotions."

[0806] In this way, this system enables effective communication among diverse users and promotes deeper understanding and positive interaction among participants.

[0807] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0808] Step 1:

[0809] Users enter their profile information using a device. Specifically, they enter data such as their name, sign language proficiency level, and areas of interest. The entered information is sent from the device to the server. In this process, a data entry interface is used, and the entered information is converted into a digital form. The output is profile data used on the server.

[0810] Step 2:

[0811] The server stores the received profile data in the database. Security protocols are applied during data storage to ensure data consistency and security. The input is the user's profile data, which forms the basis for creating records in the database. The output is a secure database record.

[0812] Step 3:

[0813] The server utilizes a generative AI model to analyze profiles stored in a database. This analysis selects the optimal user pair based on different learning levels and interests. The input is the contents of the user database, and the output is a list of pairing candidates. The AI ​​model uses machine learning algorithms to learn from past successes and improve pairing accuracy.

[0814] Step 4:

[0815] The server plans appropriate interaction activities based on paired users. Specifically, it determines the date, time, and method of the interaction event and generates an event link based on each user's schedule. The input is pairing information, and the output is detailed information about the interaction event. The server executes a scheduling algorithm to select the optimal time and method.

[0816] Step 5:

[0817] The device uses video call functionality to collect the user's facial expressions and voice in real time during interaction activities. The collected data is sent to a server. The input is the user's video call data, and the output is the data sent to the server. The device uses a high-resolution camera and microphone to acquire the data.

[0818] Step 6:

[0819] The server analyzes the transmitted facial and audio data using EmotionAI software. Through this analysis, it recognizes the user's emotions and generates appropriate feedback. The input is video call data, and the output is the result of the emotion analysis and the feedback. Specific actions include executing a facial expression analysis algorithm.

[0820] Step 7:

[0821] Based on the analysis results, the server sends feedback tailored to the user's emotions to the terminal via the communication platform. The input is the result of the emotion analysis, and the output is the user's feedback message. The server uses a messaging API to deliver the feedback.

[0822] (Application Example 2)

[0823] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0824] In online international sign language communication, a challenge is to efficiently match users with different learning levels and interests, and to enrich their interactions. Furthermore, there is a need to provide real-time sign language sessions that take users' emotions into consideration, and to enable deeper learning and understanding through visual aids.

[0825] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0826] In this invention, the server includes means for automatically matching users with different sign language learning levels and interests to facilitate international online communication; means for automatically planning and conducting interaction events for the matched users; and means for analyzing the user's facial expressions and voice, recognizing their emotions, and adjusting the interaction experience based on the results. This makes it possible to provide interactive sign language sessions that take the user's emotions into account through visual devices such as smart glasses.

[0827] "Online international communication" refers to technology that enables the exchange of information across national borders using the internet.

[0828] "Sign language learning level and interests" refers to the level of proficiency in sign language knowledge and skills, as well as the specific areas of interest of each individual user.

[0829] "Automatic matching" refers to a process where the system mechanically pairs users together based on pre-set conditions.

[0830] "Automatically planning and implementing interaction events" means that the program automatically creates and executes a communication space based on the combination of users.

[0831] "Real-time two-way translation" refers to a function that instantly translates different languages ​​during a conversation via an information processing device.

[0832] "Analyzing facial expressions and voice to recognize emotions" means using facial expression analysis software and voice recognition technology to analyze data on the user's facial expressions and voice to identify their emotions.

[0833] "Adjusting the interaction experience" refers to operations that optimize the flow of communication based on users' emotions and reactions, in order to provide a better experience.

[0834] "Receiving a sign language session through an information processing device" refers to receiving a sign language exchange conducted over a network using electronic devices.

[0835] "Displaying on a visual device" means displaying information on a vision-related device such as smart glasses or a display screen.

[0836] To realize this invention, three main elements—a server, a terminal, and a user—work together.

[0837] The server first maintains a database where users register their sign language learning level and interests. Once a user registers their profile, the server automatically matches them with other users based on this information and plans appropriate interaction events. The server also provides a real-time, two-way translation function between sign language and spoken language, facilitating smooth online communication.

[0838] The terminal primarily functions as a device for analyzing the user's emotions. Using visual devices such as smart glasses, the user's facial expressions and voice are analyzed using emotion recognition software (e.g., OpenFace), and the user's emotions are transmitted to a server. This emotion data is processed on the server, and the interaction experience is then refined.

[0839] Users can participate in sign language sessions and enjoy interactive communication with other users using smart glasses or information processing devices. Through smart glasses, they can visually receive others' sign language in real time and understand their emotions. For example, a scenario is envisioned where a sign language learner living in Japan has a real-time session with another learner in the United States to discuss environmental protection. In this case, the system can detect the user's passionate emotions, enabling a more interactive and deeper dialogue.

[0840] As an example of a prompt, you could use the request: "Create an online environment where sign language learners from around the world can learn together. Design a system that recognizes each user's emotions and learning level and takes this into consideration during sign language sessions."

[0841] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0842] Step 1:

[0843] Users register their sign language learning level and interests using a terminal. The input data is the user's profile information, which is sent to the server. The server stores this data in a database, which forms the basis for matching.

[0844] Step 2:

[0845] The server automatically matches users based on stored profile information. The input data consists of multiple user profiles and their associated interests and learning levels. The server uses an algorithm to identify users with common interests and similar learning levels and generates an output that links their profiles.

[0846] Step 3:

[0847] The server automatically plans interaction events for matched users and sends notifications to their devices. The inputs are the information of the matched users and the matching results. The server automatically generates event details and provides output to the users' devices notifying them of the schedule and how to participate.

[0848] Step 4:

[0849] The device uses visual devices such as smart glasses to collect the user's facial expressions and voice. The input consists of real-time facial expression and voice data. Emotion recognition software is used to analyze this data and generate an output that identifies the user's emotions.

[0850] Step 5:

[0851] The server receives emotional data sent from the terminal and adjusts the interaction experience. The input is the result of the emotional analysis. Based on this result, the server adjusts the progress of the communication and generates output that provides appropriate feedback to the user.

[0852] Step 6:

[0853] Users receive real-time sign language sessions and interact with other users through smart glasses. Inputs include sign language video data from other users and adjustment feedback from the server. Outputs include the user's own learning improvement and enhanced communication, which the device displays visually.

[0854] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0855] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0856] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0857] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0858] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0859] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0860] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0861] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0862] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0863] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0864] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0865] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0866] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0867] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0868] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0869] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0870] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0871] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0872] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0873] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0874] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0875] The following is further disclosed regarding the embodiments described above.

[0876] (Claim 1)

[0877] To facilitate international communication online, a means of automatically matching users with different sign language learning levels and interests is needed.

[0878] A means to automatically plan and implement interaction events for matched users,

[0879] A means of providing real-time, two-way translation between sign language and spoken language to support communication between users,

[0880] A system that includes this.

[0881] (Claim 2)

[0882] The system according to claim 1, which generates information about the schedule of a social gathering and sends it to the participants' terminals to notify them of the details of the gathering.

[0883] (Claim 3)

[0884] The system according to claim 1, which collects feedback from participants and stores data for use in improving the system.

[0885] "Example 1"

[0886] (Claim 1)

[0887] To facilitate online international dialogue, a means of automatically matching users with different learning stages and interests is needed.

[0888] A means to automatically plan and execute interaction activities for matched users,

[0889] A means of providing real-time, two-way translation between sign language and spoken language to support communication between users,

[0890] By generating information for participation in exchange events and sending it to users' information terminals, it provides a means of informing them of the details of the conversation.

[0891] A means of performing real-time translation using a generative AI model during an interaction event, and displaying and outputting the results,

[0892] A means of collecting feedback from participants and accumulating data to use for system improvement,

[0893] A system that includes this.

[0894] (Claim 2)

[0895] The system according to claim 1, which generates information for participation in an exchange event and transmits it to the user's information terminal to inform them of the details of the conversation.

[0896] (Claim 3)

[0897] The system according to claim 1, which collects opinions from participants and stores data for use in improving the system.

[0898] "Application Example 1"

[0899] (Claim 1)

[0900] To facilitate multinational information dissemination online, a means to automatically search for users with different sign language learning levels and interests,

[0901] A means to automatically plan and carry out meetings for matched users,

[0902] A means of supporting information exchange between users by performing real-time conversion between sign language and spoken language,

[0903] A means of using image input to convert sign language into speech, and converting speech information into visual information,

[0904] A system that includes this.

[0905] (Claim 2)

[0906] The system according to claim 1, which generates meeting schedule information and transmits it to participants' computers to notify them of the details of the meeting.

[0907] (Claim 3)

[0908] The system according to claim 1, which collects evaluations from participants and accumulates information that contributes to improving the system.

[0909] "Example 2 of combining an emotion engine"

[0910] (Claim 1)

[0911] To facilitate international communication online, a means of automatically pairing users with different skill learning levels and interests,

[0912] A means of automatically planning and implementing interaction activities for paired users,

[0913] A means of recognizing the user's emotions using emotion analysis functionality and generating feedback based on those emotions,

[0914] A means of providing real-time, two-way translation between voice and sign language to support communication between users,

[0915] A system that includes this.

[0916] (Claim 2)

[0917] The system according to claim 1, which generates information about the planned exchange activities and transmits it to the devices used by participants to notify them of the details of the exchange.

[0918] (Claim 3)

[0919] The system according to claim 1, which collects opinions from participants and stores data for use in improving the system.

[0920] "Application example 2 when combining with an emotional engine"

[0921] (Claim 1)

[0922] To facilitate international communication online, a means of automatically matching users with different sign language learning levels and interests is needed.

[0923] A means to automatically plan and implement interaction events for matched users,

[0924] A means of providing real-time, two-way translation between sign language and spoken language to support communication between users,

[0925] A means for analyzing the user's facial expressions and voice, recognizing emotions, and adjusting the interaction experience based on the results,

[0926] A means by which a user receives a sign language session through an information processing device and displays it on a visual device,

[0927] A system that includes this.

[0928] (Claim 2)

[0929] The system according to claim 1, which generates information about the schedule of a social gathering and sends it to the participants' terminals to notify them of the details of the gathering.

[0930] (Claim 3)

[0931] The system according to claim 1, which collects feedback from participants and stores data for use in improving the system. [Explanation of Symbols]

[0932] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. To facilitate international communication online, a means of automatically matching users with different sign language learning levels and interests, A means to automatically plan and implement interaction events for matched users, A means of providing real-time, two-way translation between sign language and spoken language to support communication between users, A system that includes this.

2. The system according to claim 1, which generates information about the schedule of a social gathering and sends it to the participants' terminals to notify them of the details of the gathering.

3. The system according to claim 1, which collects feedback from participants and stores data for use in improving the system.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A