System

A hybrid system using small-scale models on user devices and large-scale cloud models for preprocessing and output generation addresses inefficiencies and privacy issues in large-scale language models, enhancing energy efficiency and user-friendliness.

JP2026017950APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119011
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Current large-scale language models consume significant computational resources, lack energy efficiency, and fail to provide adequate privacy protection, leading to inefficient and user-unfriendly interactions, especially for new users.

Method used

A system combining small-scale language models on user devices for preprocessing and large-scale models on the cloud for high-quality output generation, with preprocessing steps that mask privacy information and clarify ambiguous queries.

Benefits of technology

This approach improves energy efficiency and privacy protection, enabling high-quality responses to user queries while accommodating a wide range of users, from individuals to companies and public institutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017950000001_ABST
    Figure 2026017950000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for parsing and preprocessing a user's input using a small language model; means for sending the preprocessed data to a large language model on a cloud; means for receiving a high quality output generated by the large language model; and means for displaying the received high quality output to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, advances in artificial intelligence technology have led to the widespread adoption of various applications using language models. However, current large-scale language models consume large amounts of computational resources and suffer from low energy efficiency. Furthermore, when receiving direct user input, privacy protection is often insufficient, leading to repeated unnecessary questions. This makes the system difficult to use, especially for users new to AI. There is a need for a system that can solve these issues, improve energy efficiency, and provide privacy protection and a user-friendly interface. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system that includes a means for analyzing and preprocessing user input using a small-scale language model, a means for transmitting the preprocessed data to a large-scale language model on the cloud, a means for receiving high-quality output generated by the large-scale language model, and a means for displaying the received high-quality output to the user. The preprocessing process masks the user's privacy information and converts the question into an appropriate format to prevent unnecessary repetition of questions. Furthermore, energy efficiency is improved by installing the small-scale language model that performs preprocessing on the device and the large-scale language model in a cloud environment. This enables the system to accommodate a wide range of users, from individuals to companies and public institutions, and realize sustainable AI use.

[0006] A "small-scale language model" is a natural language processing model that can be executed using small, energy-efficient capacity and computational resources installed on a device.

[0007] "Preprocessing" is a process that analyzes user input and performs optimizations such as eliminating ambiguous expressions and masking privacy information.

[0008] A "large-scale language model" is a high-performance natural language processing model trained using large amounts of computing resources and data.

[0009] "On the cloud" refers to a server environment that can be accessed remotely via the Internet.

[0010] "High-quality output" is an accurate and detailed response to a user's request, generated by a large language model.

[0011] "Masking privacy information" means protecting a user's personal information and location information by converting them so that they cannot be identified.

[0012] "Converting to an appropriate format" means shaping the input question into a format that is efficient and reduces the risk of errors in the answer.

[0013] "Improved energy efficiency" means reducing the amount of power required to perform the same task.

[0014] "Sustainable use of AI" means utilizing AI technology over the long term while minimizing the environmental impact.

[0015] "Displaying to the user" means visually outputting the processing results on the screen of the user's device.

[0016] "Preventing unnecessary repetition of questions" means preventing the need to process the same or similar questions multiple times. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention is a system that combines small-scale language models (hereinafter referred to as small-scale LLMs) and large-scale language models (hereinafter referred to as large-scale LLMs), which enables efficient processing of user input and provides high-quality output. This system aims to protect privacy and improve energy efficiency, among other things.

[0039] System configuration

[0040] This system is intended to be used by users through devices such as smartphones and tablets. The system includes the following components:

[0041] Terminal

[0042] Small scale LLM

[0043] Data reception and transmission module

[0044] Display Module

[0045] server

[0046] large scale LLM

[0047] Data reception and transmission module

[0048] Information Collection Module

[0049] System Operation

[0050] When a user uses the system to ask a question, the following process is carried out.

[0051] Capturing User Input

[0052] A user inputs a question into a smartphone app. For example, they input a question like, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[0053] Pretreatment by small-scale LLM

[0054] The device first sends the captured input to a small-scale LLM for preprocessing. Specifically, it removes ambiguous expressions from the question and masks the user's privacy information (e.g., current location information). For example, it clarifies the expression "nearby" to "within a 2km radius." This preprocessing aims to clarify the question while protecting the user's privacy.

[0055] Sending preprocessed data

[0056] The pre-processed data is sent to the server by the terminal's data transmission module, which receives the data and passes it to the large-scale LLM.

[0057] Generating high-quality output with large-scale LLM

[0058] The server-side large-scale LLM analyzes the preprocessed queries and generates high-quality output that meets the user's requirements. For example, it generates a list of "highly rated restaurants within a 2km radius." The large-scale LLM collects and organizes information using databases and external APIs as needed.

[0059] Send and view high-quality output

[0060] The generated high-quality output is sent to the terminal by the server's data transmission module, and the terminal displays the received data to the user using the display module, allowing the user to accurately and quickly obtain the information they desire while protecting their privacy.

[0061] Specific examples

[0062] As a specific example, the following scenario is assumed.

[0063] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[0064] The terminal receives the input and passes it to the small-scale LLM.

[0065] 2. A small LLM preprocesses the user input.

[0066] For example, the vague expression "nearby" is made more specific, and the location information is masked to the city / ward / town / village level.

[0067] 3. The preprocessed data is sent to the server.

[0068] The data transmission module of the terminal transmits the preprocessed data to the server.

[0069] 4. The server's large-scale LLM generates high-quality restaurant lists.

[0070] The server uses a database or external API to retrieve information about nearby highly rated restaurants.

[0071] 5. The generated list is sent to the terminal and displayed to the user.

[0072] The terminal receives the list and visually presents it to the user.

[0073] In this way, the system provides high-quality, energy-efficient information while protecting user privacy.

[0074] The processing flow will be explained below.

[0075] Step 1:

[0076] A user types a question into a smartphone app: "What are some recommended restaurants nearby?"

[0077] Step 2:

[0078] The terminal receives the user's input at a data receiving module.

[0079] Step 3:

[0080] The device passes the received input to the small-scale LLM, which analyzes the question and translates it into a more specific expression. For example, it converts "nearby" to "within a 2km radius" and masks the user's current location to the city level.

[0081] Step 4:

[0082] The terminal transmits the pre-processed data to the server via a data transmission module.

[0083] Step 5:

[0084] The server receives the preprocessed data transmitted from the terminal at a data receiving module.

[0085] Step 6:

[0086] The server passes the received data to a large-scale LLM, which analyzes the user's question and generates high-quality output. For example, it may use a database or external API to collect information on highly rated restaurants within a 2km radius.

[0087] Step 7:

[0088] The server transmits high-quality output to the terminal via a data transmission module.

[0089] Step 8:

[0090] The terminal receives the high-quality output transmitted from the server at a data receiving module.

[0091] Step 9:

[0092] The data received by the terminal is visually displayed to the user in a display module, for example, a list of highly rated restaurants is displayed on the screen.

[0093] This series of processes allows users to quickly obtain high-quality information that protects their privacy in an energy-efficient manner.

[0094] Example 1

[0095] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0096] Modern information processing systems are required to quickly and accurately provide high-quality output in response to user-entered questions. However, achieving this goal poses several challenges. First, from the perspective of privacy protection, users' personal information must be handled securely. Second, if the input question is ambiguous, preprocessing is required to make it more specific. Furthermore, generating high-quality output requires processing large amounts of data, and improving energy efficiency for this purpose is also an important challenge.

[0097] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0098] In this invention, the server includes means for analyzing and preprocessing user input using a small-scale language model installed on the user terminal, means for transmitting the preprocessed data to a large-scale language model on the cloud, and means for receiving high-quality output generated by the large-scale language model, thereby enabling ambiguous questions to be specified and providing high-quality information with high energy efficiency while protecting user privacy.

[0099] A "small language model" is a language processing algorithm that runs on a terminal and analyzes and preprocesses user input.

[0100] "Preprocessing" refers to the process of removing ambiguous expressions from input data and anonymizing user privacy information as necessary.

[0101] "Anonymization" refers to making a user's information in a state where it is not possible to identify the individual in order to protect the user's privacy information.

[0102] "Concreteization" is the process of converting vague expressions into clear, specific expressions.

[0103] A "large-scale language model" is an advanced language processing algorithm that runs on the cloud and produces high-quality output.

[0104] "High-quality output" refers to relevant and accurate answers and information to users' questions.

[0105] A "database" is a system for systematically storing and managing data.

[0106] An "external API" is a program interface that connects to external databases and services to obtain information.

[0107] A "secure communication protocol" is a communication rule for safely sending and receiving data.

[0108] A "user terminal" is a computer device that can be directly operated by a user, such as a smartphone or tablet.

[0109] "Analysis" is the act of breaking down and examining input data in order to understand it and perform the necessary processing.

[0110] "Masking" is the act of hiding or making certain information invisible.

[0111] The present invention is an information processing system that combines a small-scale language model installed on a user terminal with a large-scale language model on the cloud. The system aims to provide high-quality output in response to user input. Specifically, the small-scale language model preprocesses data entered by the user, and the preprocessed data is sent to a large-scale language model on the cloud for analysis. The system then receives the high-quality output generated by the large-scale language model and displays it to the user.

[0112] Hardware and software used

[0113] User device: A computing device that can be directly operated by a user, such as a smartphone or tablet. The device is equipped with a small language model, a data reception / transmission module, and a display module.

[0114] Cloud server: A server on which a large-scale language model runs, and includes a data reception / transmission module and an information collection module.

[0115] Small language model behavior

[0116] A small language model installed on the user's device analyzes data entered by the user (e.g., "What are some recommended restaurants nearby?") and performs preprocessing. This preprocessing involves concretizing ambiguous expressions and anonymizing private information. For example, the ambiguous expression "nearby" is concretized as "within a 2km radius," and location information is converted to the city, town, or village level.

[0117] Sending data

[0118] The pre-processed data is then sent to the cloud server via the device's data transmission module, where data security is ensured using secure communication protocols such as HTTPS.

[0119] How large language models work

[0120] The preprocessed data sent to the cloud server is analyzed by a large-scale language model, which gathers information from databases and external APIs to generate the best answer to the user's question, such as a list of "highly rated restaurants within a 2km radius."

[0121] Send and view high-quality output

[0122] The generated high-quality output is sent to the user terminal by the server's data transmission module, and the terminal visually presents the received data to the user via the display module, allowing the user to obtain the information they need accurately and quickly while protecting their privacy.

[0123] Specific examples

[0124] 1. User Input: A user types into their smartphone, "What are some recommended restaurants nearby?"

[0125] 2. Preprocessing: A small language model concretizes the expression "nearby" to "within a 2km radius" and anonymizes location information to the city / ward / town / village level.

[0126] 3. Data transmission: The preprocessed data is transmitted to the cloud server using HTTPS.

[0127] 4. Parsing and output generation: A large language model uses databases and external APIs to create a list of "top rated restaurants within a 2km radius."

[0128] 5. Sending and displaying the output: The generated list is sent to the user terminal and presented to the user via the display module.

[0129] In this way, the present invention provides high-quality information while protecting the user's privacy.

[0130] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0131] Step 1: Capturing User Input

[0132] A user types "What are some recommended restaurants nearby?" into a smartphone app. The device captures this input in a data receiving module and passes it to a small language model as text data. At this stage, the input data is the raw text entered by the user.

[0133] Step 2: Preprocessing with small-scale LLM

[0134] A small language model on the device receives the captured input text and analyzes it. For the input text "What are some recommended restaurants nearby?", the ambiguous "nearby" is made more specific to "within a 2km radius," and the user's location information is further anonymized at the city / ward / town / village level. The output of this step is the preprocessed text data "What are some recommended restaurants within a 2km radius?"

[0135] Step 3: Sending preprocessed data

[0136] The preprocessed data is received by the terminal's data transmission module and sent to the cloud server. HTTPS is used for this transmission to ensure data security. The input is the preprocessed text data, and the output is the encoded data transmission.

[0137] Step 4: Receiving data on the server

[0138] The server receives data sent from the device through a data reception module, decodes the received encoded data, and prepares it for passing to a large-scale language model, where the input is the encoded preprocessed data and the output is the decoded preprocessed data.

[0139] Step 5: Generating high-quality output with large-scale LLM

[0140] The server's large-scale language model analyzes the preprocessed data. Specifically, for the question "What are some recommended restaurants within a 2km radius?", it collects information from databases and external APIs and generates a list of "highly rated restaurants within a 2km radius." The input in this step is the decoded preprocessed data, and the output is a high-quality restaurant list.

[0141] Step 6: Send high-quality output

[0142] The generated high-quality output is encoded by the server's data transmission module and sent to the terminal. Again, HTTPS is used to transmit data securely. The input is a high-quality restaurant list, and the output is the encoded restaurant list.

[0143] Step 7: Receive and display data on your device

[0144] The terminal's data receiving module receives high-quality output from the server. The received data is decoded and passed to the display module. The display module displays the information in a visually easy-to-understand format for the user, for example, as a "list of highly rated restaurants." The input in this step is the encoded restaurant list, and the output is restaurant information visually displayed to the user.

[0145] This allows users to obtain the information they need accurately and quickly while their privacy is protected.

[0146] (Application example 1)

[0147] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0148] Conventional language modeling systems have problems such as insufficient protection of user privacy information and difficulty in obtaining high-quality answers due to ambiguous user questions. Energy efficiency is also an issue. Food delivery services, in particular, require users to quickly and accurately find nearby restaurant recommendations.

[0149] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0150] In this invention, the server includes: means for analyzing and preprocessing user input using a small-scale language model; means for sending the preprocessed data to a large-scale language model on the cloud; means for receiving high-quality output generated by the large-scale language model; means for displaying the received high-quality output to the user; means for converting ambiguous expressions into specific conditions and anonymizing location information in the preprocessing; means for masking privacy information to the city / ward / town / village level to protect user privacy; means for generating a restaurant list in descending order of rating; means for collecting and organizing information based on the preprocessed data; and means for generating high-quality output from prompt sentences using a generative AI model. This enables ambiguous questions to be converted into specific, high-quality answers while protecting user privacy. It also enables users to quickly and accurately find recommended nearby restaurants for food delivery.

[0151] A "small language model" is a relatively small language processing system capable of analyzing user input and converting vague expressions into concrete terms.

[0152] "Preprocessing" is the process of analyzing user input information, concretizing ambiguous expressions, and anonymizing and masking privacy information.

[0153] A "large-scale language model in the cloud" is a large-scale language processing system accessed via the internet for complex analysis and high-quality output generation.

[0154] "High-quality output" is accurate and detailed information generated by a large-scale language model and tailored to the user's needs.

[0155] "Means for displaying to a user" refers to a display or display module for visually presenting the received output on a terminal.

[0156] "Ambiguous expressions" refer to parts of the user's input that lack specificity or are unclear.

[0157] "Specific conditions" are those that convert vague expressions into clear and detailed information.

[0158] "Means for anonymizing location information" refers to a method for protecting privacy by deleting or converting the user's current location to a level that makes it identifiable.

[0159] "Means for masking privacy information to the city, ward, town, or village level" refers to a method for preventing individuals from being identified by limiting the user's location information to the city, ward, town, or village level.

[0160] The "means for generating a restaurant list in descending order of rating" is a method for sorting nearby restaurants in a food delivery service by rating criteria.

[0161] "Means for collecting and organizing information based on preprocessed data" refers to a method for obtaining necessary information from preprocessed data and organizing it appropriately.

[0162] A "generative AI model" is an artificial intelligence language processing model that provides appropriate output based on an input prompt.

[0163] A "prompt" is an instruction or question given to an AI model to generate a specific output.

[0164] The system for implementing this invention consists of a terminal that analyzes and preprocesses user input using a small-scale language model, and a server that generates high-quality output using a large-scale language model on the cloud.

[0165] System configuration

[0166] The system consists of the following hardware and software:

[0167] Hardware

[0168] User devices (smartphones, tablets)

[0169] Cloud server with graphics card

[0170] software

[0171] Small language model (runs on device)

[0172] Large-scale language model (running on cloud servers)

[0173] Transformers Library (Hugging Face)

[0174] requests library (data communication)

[0175] React Native (user interface)

[0176] System Operation

[0177] The system operates in the following manner to efficiently process user input information and provide high quality output.

[0178] 1. Capturing User Input

[0179] A user uses a smartphone app to input a question, for example, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[0180] 2. Pretreatment by small-scale LLM

[0181] User input is first sent to a small language model on the device for preprocessing. For example, a vague expression like "nearby" is converted into a more specific condition like "within a 2km radius." Additionally, user location information is anonymized at the city / ward / town / village level, protecting user privacy.

[0182] 3. Sending preprocessed data

[0183] After the preprocessing is complete, the data is sent to the cloud server by the terminal's data transmission module, which receives the data and passes it to a large-scale language model.

[0184] 4. Generating high-quality output using large-scale LLM

[0185] A large-scale language model on a cloud server analyzes the preprocessed questions and generates high-quality output, such as a list of "top-rated restaurants within a 2km radius." The large-scale language model collects and organizes information using databases and external APIs as needed.

[0186] 5. Send and display high-quality output

[0187] The generated high-quality output is sent to the device by the cloud server's data transmission module, and the device displays the received data to the user via the display module, allowing the user to accurately and quickly obtain the information they desire while protecting their privacy.

[0188] Examples of concrete examples and prompts

[0189] A concrete example of how the system has been implemented is a food delivery application, which works as follows:

[0190] 1. A user types into the app, "What are some recommended restaurants nearby?"

[0191] 2. The app converts the input into "What are some recommended restaurants within a 2km radius?" and anonymizes the location information to "Shinjuku-ku, Tokyo" before sending it to the server.

[0192] 3. The server uses an AI model to generate a list of highly rated restaurants nearby.

[0193] 4. Send the list to the app and display it to the user.

[0194] Example prompt sentence:

[0195] Please list recommended restaurants within a 2km radius. The user's anonymized location is "Shinjuku-ku, Tokyo."

[0196] In this way, it is possible to provide highly accurate restaurant information for food delivery while ensuring the user's privacy.

[0197] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0198] Step 1:

[0199] A user opens a smartphone application and inputs the question, "What are some recommended restaurants nearby?" This input is captured on the user's device and passed to a data receiving module. The input data is natural language text, and the operation to capture the input data is performed via a user interface. In this case, the input is "What are some recommended restaurants nearby?" and the output is the user's input text data.

[0200] Step 2:

[0201] The device passes the captured user input to a small-scale language model and begins preprocessing. Specifically, the small-scale LLM concretizes the ambiguous expression "nearby" to "within a 2km radius" and anonymizes the user's location information to "Shinjuku Ward, Tokyo" at the city / ward / town / village level. The input for preprocessing is the captured user's natural language text and location information, and the output is a preprocessed concretized question: "Recommend restaurants within a 2km radius (Shinjuku Ward, Tokyo)."

[0202] Step 3:

[0203] The preprocessed data is sent to the cloud server by the terminal's data transmission module. During the transmission process, the data is encrypted using a protocol such as HTTPS request and sent to the server. The input is the preprocessed query data, and the output is the completion of data transmission to the cloud server.

[0204] Step 4:

[0205] The cloud server passes the received preprocessed data to a large-scale language model (generative AI model). The large-scale LLM uses external databases and APIs to provide high-quality information. Specifically, it generates a list of "highly rated restaurants within a 2km radius." The input of the large-scale LLM is the preprocessed, specified question data, and the output is a list of highly rated restaurants.

[0206] Step 5:

[0207] The server receives the high-quality output (restaurant list) generated by the large-scale LLM and sends it to the terminal. The transmission process is also encrypted and is carried out using protocols such as HTTPS requests. The input here is the generated restaurant list data, and the output is the completion of transmission of the restaurant list to the terminal.

[0208] Step 6:

[0209] The terminal passes the restaurant list received from the server to the display module, which then visually presents it to the user. Specifically, the restaurant list is displayed on the application interface using a framework such as React Native. The input here is the received restaurant list data, and the output is the specific restaurant list visually displayed to the user.

[0210] This series of processing steps allows users to quickly obtain high-quality restaurant information for food delivery services while protecting their privacy.

[0211] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0212] This invention is a system that combines a small-scale language model (hereinafter referred to as small-scale LLM) with a large-scale language model (hereinafter referred to as large-scale LLM), and adds an emotion engine that recognizes user emotions. It efficiently processes user input and provides high-quality output according to the user's emotions. The system aims to protect privacy, improve energy efficiency, and respond to user emotions.

[0213] System configuration

[0214] The system is accessed by users through specific input devices (such as smartphones or tablets) and includes the following components:

[0215] Terminal

[0216] Small scale LLM

[0217] Emotion Engine

[0218] Data reception and transmission module

[0219] Display Module

[0220] server

[0221] large scale LLM

[0222] Data reception and transmission module

[0223] Information Collection Module

[0224] System Operation

[0225] When a user uses the system to ask a question, the following process is carried out.

[0226] Capturing User Input

[0227] A user inputs a question into a smartphone app. For example, they input a question like, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[0228] Emotion recognition with emotion engine

[0229] The device sends the captured input to an emotion engine to recognize the user's emotions, for example, determining whether the user is in a hurry or relaxed state based on the style and tone of the question.

[0230] Pretreatment by small-scale LLM

[0231] The device uses the emotion information obtained from the emotion engine and passes it to the small-scale LLM. The small-scale LLM analyzes the question and converts it into a more specific expression. For example, it converts the vague expression "nearby" into "within a 2km radius" and masks the user's current location information to the city / ward / town / village level. If the device detects a sense of urgency, it sets a priority to speed up processing.

[0232] Sending preprocessed data

[0233] The pre-processed data is sent to the server by the terminal's data transmission module, which receives the data and passes it to the large-scale LLM.

[0234] Generating high-quality output with large-scale LLM

[0235] The server-side large-scale LLM analyzes the preprocessed queries and generates high-quality output that meets the user's requirements. For example, it generates a list of "highly rated restaurants within a 2km radius." The large-scale LLM collects and organizes information using databases and external APIs as needed.

[0236] Send and view high-quality output

[0237] The generated high-quality output is transmitted to the terminal by the server's data transmission module, and the terminal displays the received data to the user using the display module. This allows the user to quickly obtain high-quality information according to their emotions while protecting their privacy.

[0238] Specific examples

[0239] As a specific example, the following scenario is assumed.

[0240] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[0241] The device receives the input and passes it to the emotion engine.

[0242] 2. The emotion engine recognizes the user's emotions.

[0243] For example, it is determined that the user is in a hurry.

[0244] 3. A small-scale LLM performs preprocessing based on emotion information.

[0245] It clarifies the vague term "nearby" by masking location information to city level, and sets priorities for faster processing based on feelings of urgency.

[0246] 4. The preprocessed data is sent to the server.

[0247] The data transmission module of the terminal transmits the preprocessed data to the server.

[0248] 5. The server's large-scale LLM generates high-quality restaurant lists.

[0249] The server uses a database or external API to collect information about nearby highly rated restaurants.

[0250] 6. The generated list is sent to the terminal and displayed to the user.

[0251] The terminal receives the list and visually presents it to the user.

[0252] In this way, the system understands users' emotions while providing high-quality information with improved energy efficiency and privacy protection.

[0253] The processing flow will be explained below.

[0254] Step 1:

[0255] A user types a question into a smartphone app: "What are some recommended restaurants nearby?"

[0256] Step 2:

[0257] The terminal receives the user's input at a data receiving module.

[0258] Step 3:

[0259] The device sends the received input to an emotion engine that analyzes the user's emotions, such as identifying emotions like "hurried" or "relaxed" based on the user's writing style and tone.

[0260] Step 4:

[0261] If the device receives the analysis results from the emotion engine and determines that the user is in a hurry, it sends the input data to a small-scale LLM and instructs it to perform high-priority processing.

[0262] Step 5:

[0263] A small LLM analyzes user input and converts vague expressions into specific ones, for example, changing "nearby" to "within a 2km radius" and masking the user's current location to city level.

[0264] Step 6:

[0265] The terminal transmits the pre-processed data to the server via a data transmission module.

[0266] Step 7:

[0267] The server receives the preprocessed data transmitted from the terminal at a data receiving module.

[0268] Step 8:

[0269] The server passes the received data to a large-scale LLM, which analyzes the data and generates high-quality output based on the user's sentiment and requirements. For example, it creates a list of highly rated restaurants within a 2km radius.

[0270] Step 9:

[0271] The high-quality output generated by the server is transmitted to the terminal by a data transmission module.

[0272] Step 10:

[0273] The terminal receives the high-quality output transmitted from the server at a data receiving module.

[0274] Step 11:

[0275] The data received by the terminal is visually displayed to the user in a display module, for example, a list of highly rated restaurants is displayed on the screen.

[0276] Through this series of steps, users can quickly obtain high-quality information that corresponds to their emotions while their privacy is protected.

[0277] Example 2

[0278] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0279] Current language model systems perform uniform processing without considering the user's emotions, resulting in a poor user experience. They also sometimes fail to adequately protect the user's privacy and can be slow to respond. Furthermore, they lack the ability to properly concretize ambiguous input, making it difficult to provide the accurate information the user desires.

[0280] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for capturing user input, means for analyzing and preprocessing the user input using a small-scale language model, means for performing sentiment analysis, means for transmitting the preprocessed data to a large-scale language model on the cloud, means for receiving high-quality output generated by the large-scale language model, means for displaying the received high-quality output to the user, and means for concretizing ambiguous expressions based on the user input and setting priorities if necessary. This enables the provision of high-quality information that takes user emotions into consideration, protection of privacy, and highly efficient data processing.

[0281] A "small language model" is a computer program that analyzes and preprocesses user input data, optimized to operate in environments with limited computing resources.

[0282] A "large-scale language model" is a computer program that uses large amounts of data and computing resources to perform advanced natural language processing and generate high-quality output.

[0283] "Preprocessing" refers to a series of operations that analyze data entered by a user and perform masking to concretize ambiguous expressions and protect privacy.

[0284] "Receiving" means receiving data sent from a sender.

[0285] "Emotion analysis" is the process of inferring a user's emotions and intentions based on the data they input.

[0286] "Data transmission" means sending pre-processed data to a designated destination.

[0287] "High-quality output" is information generated by a large-scale language model to accurately respond to user requests.

[0288] A "data receiving module" is hardware or software that has the function of receiving data from the outside and passing it on to internal processing.

[0289] A "data transmission module" is hardware or software that has the function of transmitting pre-processed or generated data to the outside.

[0290] A "display module" is hardware or software that has the function of visually displaying received information to a user.

[0291] This system combines small-scale and large-scale language models and adds an emotion analysis engine to recognize user emotions. It efficiently processes user input and provides high-quality output that reflects the user's emotions. This system aims to protect privacy, improve energy efficiency, and respond to user emotions.

[0292] System configuration

[0293] The system is accessed by users through specific input devices (such as smartphones or tablets) and includes the following components:

[0294] Terminal

[0295] Small language models

[0296] It is a computer program that analyzes and preprocesses user input data.

[0297] Sentiment Analysis Engine

[0298] It is a computer program that analyzes emotions and intentions based on user input data.

[0299] Data Receiving Module

[0300] It has the function of receiving input from the user and passing it on to internal processing.

[0301] Data Transmission Module

[0302] It has a function for sending preprocessed data to the server.

[0303] Display Module

[0304] It has the ability to visually display high quality output received from the server to the user.

[0305] server

[0306] Large-scale language models

[0307] It is a computer program that performs advanced natural language processing and generates high-quality output that meets user requirements.

[0308] Data Receiving Module

[0309] It has the function of receiving data sent from the terminal and passing it on to internal processing.

[0310] Information Collection Module

[0311] It has the ability to collect information from databases and external APIs as needed and organize the data.

[0312] Specific examples

[0313] When a user asks a question

[0314] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[0315] The terminal's data receiving module captures this input.

[0316] 2. The sentiment analysis engine recognizes the user's emotions.

[0317] For example, the style and tone of the question can determine whether the user is in a hurry or relaxed.

[0318] 3. A small language model performs preprocessing based on emotion information.

[0319] The vague expression "nearby" is specified as "within a 2km radius," location information is masked to the city level, and priority is set for fast processing based on the feeling of urgency.

[0320] 4. The preprocessed data is sent to the server.

[0321] The data transmission module of the terminal transmits the preprocessed data to the server.

[0322] 5. The server's large-scale language model generates high-quality restaurant lists.

[0323] A large language model uses databases and external APIs to create a list of "highly rated restaurants within a 2km radius."

[0324] 6. The generated list is sent to the terminal and displayed to the user.

[0325] The terminal receives the list and visually presents it to the user using a display module.

[0326] Examples of prompt statements

[0327] For example, if a user inputs a question such as "What are some recommended restaurants nearby?", the above process will promptly provide high-quality information while taking into consideration the user's feelings.

[0328] In this way, this system understands the user's emotions and provides high-quality information while protecting privacy.

[0329] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0330] Step 1: Capturing User Input

[0331] A user types into their smartphone, "What are some recommended restaurants near me?"

[0332] The terminal's data reception module captures this input. Specifically, it acquires the string entered in the text box as internal data.

[0333] Input: User question (e.g., "What are some recommended restaurants nearby?")

[0334] Output: Captured user question data (e.g., "What are some recommended restaurants nearby?")

[0335] Step 2: Emotion recognition by the emotion engine

[0336] The device sends the captured user question to the emotion engine, which analyzes emotions from the input text.

[0337] Input: User question data (e.g., "What are some recommended restaurants nearby?")

[0338] Output: Emotional information (e.g., user is in a hurry)

[0339] Step 3: Preprocessing with small-scale LLM

[0340] The device passes the captured question data based on the emotion information to the small-scale LLM, which analyzes the question and makes ambiguous expressions concrete.

[0341] Input: User question data and emotion information (e.g., "What are some recommended restaurants nearby?", emotion: "I'm in a hurry")

[0342] Specific behavior: Converts the vague expression "nearby" to "within a 2km radius," masks location information to the city / ward / town / village level for privacy reasons, and sets priorities for faster processing.

[0343] Output: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", sentiment: in a hurry)

[0344] Step 4: Sending preprocessed data

[0345] The data transmission module of the terminal transmits the preprocessed data to the server, specifically, by using an HTTP request.

[0346] Input: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", Sentiment: Hurry)

[0347] Output: Data sent to the server

[0348] Step 5: Generating high-quality output with large-scale LLM

[0349] The server receives the pre-processed data and passes it to the large-scale LLM, which uses databases and external APIs to generate high-quality output.

[0350] Input: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", Sentiment: Hurry)

[0351] Specific behavior: Performs database queries and external API calls, and generates a list of restaurants based on the obtained data.

[0352] Output: High-quality output data (e.g., "List of highly rated restaurants within a 2km radius")

[0353] Step 6: Send and view high-quality output

[0354] The data transmission module of the server transmits the generated high-quality output to the terminal.

[0355] The terminal receives the output sent from the server and visually displays it to the user using a display module.

[0356] Input: High-quality output data (e.g., "List of highly rated restaurants within a 2km radius")

[0357] Specific behavior: Renders the received data in the UI and displays it to the user.

[0358] Output: High-quality information visually presented to the user

[0359] In this way, the system understands users' emotions while providing high-quality information with improved energy efficiency and privacy protection.

[0360] (Application example 2)

[0361] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0362] Conventional systems have had difficulty in accurately recognizing emotions in response to user input and providing personalized, high-quality information based on that emotion. Furthermore, while there is a need to efficiently provide information tailored to the user's emotions, there are insufficient measures to protect privacy and prevent unnecessary repetition of questions. Understanding customer emotions and providing real-time services based on those emotions is particularly important in brick-and-mortar stores, but conventional technologies have had difficulty meeting this requirement.

[0363] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing and preprocessing a user's input using a small-scale language model, means for transmitting the preprocessed data to a large-scale language model on the cloud, means for receiving high-quality output generated by the large-scale language model, means for displaying the received high-quality output to the user, means for recognizing emotions from the user's input, and means for adjusting data preprocessing based on the recognized emotions. This makes it possible to provide information according to the user's emotions in real time, protect privacy, and prevent unnecessary repetition of questions.

[0364] A "small language model" is a lightweight, energy-efficient natural language processing model used to parse and preprocess user input.

[0365] A "large-scale language model" is a high-performance natural language processing model that runs on the cloud and generates high-quality output based on pre-processed data.

[0366] "User input" refers to data such as text information and voice information that a user provides to the system.

[0367] "Preprocessing" refers to the initial processing of user input to parse it and convert it into a format that meets specific needs.

[0368] "Sending to the cloud" refers to transferring data from the user's terminal to a remote server.

[0369] "High-quality output" refers to highly accurate information and results generated by a large-scale language model that are tailored to the user's needs.

[0370] "Receiving" refers to receiving the generated output from the server at the user's terminal.

[0371] "Displaying" refers to visually presenting the received high quality output to a user.

[0372] "Emotion recognition" refers to the process of determining a user's emotional state from their input.

[0373] "Adjusting data preprocessing" refers to changing and optimizing the content and methods of preprocessing based on the recognized emotions.

[0374] The system for implementing this invention analyzes user input, recognizes emotions, and provides high-quality output based on those emotions. The system consists of two main components: a terminal and a server.

[0375] Program Overview

[0376] 1. Capturing user input

[0377] The user inputs data through a terminal (e.g., a smartphone). The input data can be text data or voice data.

[0378] 2. Emotional Recognition

[0379] The device passes the input data to an emotion engine (e.g., Hugging Face sentiment-analysis) to recognize the user's emotion. For example, if the user input is "Are there any deals?", the device recognizes the user's emotion as "happy" based on that input.

[0380] 3. Pretreatment

[0381] It uses a small language model (e.g., T5-small) to analyze user input and convert it into a concrete form while taking into account emotional information, for example, converting the vague expression "great deal" into "today's sale items," and adjusts the processing speed based on the emotion.

[0382] 4. Data transmission

[0383] The pre-processed data is transmitted to the server by the data transmission module of the terminal.

[0384] 5. Processing with large-scale language models

[0385] The server uses large-scale language models to analyze the pre-processed data and generate high-quality output, such as a list of today's sale items.

[0386] 6. Send and display high-quality output

[0387] The server generates high-quality output and sends it to the terminal, which visually displays the received data to the user.

[0388] Hardware and software used

[0389] Device: User device such as a smartphone or tablet

[0390] Emotion Engine: A sentiment-analysis model of Hugging Face

[0391] Small language models: Natural language processing models such as T5-small

[0392] Server: High-performance server on the cloud

[0393] Large-scale language model: High-performance natural language processing model running on a server

[0394] Specific examples

[0395] The specific scenario is as follows:

[0396] 1. The user speaks into the smartphone app and says, "Do you have any deals?"

[0397] 2. The device captures the input and recognizes the "happy" emotion using its emotion engine.

[0398] 3. Small LLMs concretize the expression "good deal" and transform it into "What's on sale today?"

[0399] 4. Send it to the server, and the large-scale LLM generates a "list of today's sale items" that meets your request.

[0400] 5. Once the list is returned, the device displays it to the user along with "Today's Sale Items."

[0401] Example prompt sentence:

[0402] Q: Are there any great deals? I'd be happy to provide information at a speed that matches my needs. A:

[0403] In this way, the system can provide personalized information based on user sentiment in real time, improving the customer experience.

[0404] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0405] Step 1:

[0406] The user types "Do you have any deals?" into the smartphone app. The input can be text or voice. The device captures this data and stores it as input data.

[0407] Step 2:

[0408] The device passes the captured input data to the emotion engine, which (Hugging Face's sentiment-analysis model) recognizes the emotion from the user's input. In this step, based on the input data (e.g., "Are there any deals available?"), the emotion engine outputs the emotion information "happy."

[0409] Step 3:

[0410] The device takes the recognized emotional information into consideration and passes the user's input to a small-scale LLM (T5-small model). The small-scale LLM analyzes the input and performs data preprocessing based on the emotional information. In this step, the device outputs preprocessed data (e.g., "What are the sale items today?") based on the input data and emotional information (e.g., "Are there any deals?" and "I'm happy").

[0411] Step 4:

[0412] The data transmission module of the device sends the preprocessed data to a server on the cloud, where the preprocessed data ("What are the sale items today?") is transferred to the server as input data.

[0413] Step 5:

[0414] The server passes the received preprocessed data to the large-scale LLM, which analyzes the preprocessed data and generates the corresponding high-quality output. In this step, the server outputs a high-quality output ("List of sale items") based on the preprocessed input data ("What items are on sale today?").

[0415] Step 6:

[0416] The server's data transmission module transmits the generated high-quality output to the terminal, where the generated data ("list of sale items") is transmitted to the terminal.

[0417] Step 7:

[0418] The terminal uses a display module to display the received high-quality data to the user. In this step, the terminal visually presents the high-quality output data ("list of sale items") to the user.

[0419] The above is the specific process flow from user input to the provision of high-quality information. Data is input and output at each step, and the next step is executed based on the results, thereby realizing the provision of highly accurate, emotion-responsive information.

[0420] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0421] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0422] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0423] [Second embodiment]

[0424] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0425] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0426] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0427] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0428] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0429] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0430] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0431] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0432] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0433] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0434] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0435] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0436] The present invention is a system that combines small-scale language models (hereinafter referred to as small-scale LLMs) and large-scale language models (hereinafter referred to as large-scale LLMs), which enables efficient processing of user input and provides high-quality output. This system aims to protect privacy and improve energy efficiency, among other things.

[0437] System configuration

[0438] This system is intended to be used by users through devices such as smartphones and tablets. The system includes the following components:

[0439] Terminal

[0440] Small scale LLM

[0441] Data reception and transmission module

[0442] Display Module

[0443] server

[0444] large scale LLM

[0445] Data reception and transmission module

[0446] Information Collection Module

[0447] System Operation

[0448] When a user uses the system to ask a question, the following process is carried out.

[0449] Capturing User Input

[0450] A user inputs a question into a smartphone app. For example, they input a question like, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[0451] Pretreatment by small-scale LLM

[0452] The device first sends the captured input to a small-scale LLM for preprocessing. Specifically, it removes ambiguous expressions from the question and masks the user's privacy information (e.g., current location information). For example, it clarifies the expression "nearby" to "within a 2km radius." This preprocessing aims to clarify the question while protecting the user's privacy.

[0453] Sending preprocessed data

[0454] The pre-processed data is sent to the server by the terminal's data transmission module, which receives the data and passes it to the large-scale LLM.

[0455] Generating high-quality output with large-scale LLM

[0456] The server-side large-scale LLM analyzes the preprocessed queries and generates high-quality output that meets the user's requirements. For example, it generates a list of "highly rated restaurants within a 2km radius." The large-scale LLM collects and organizes information using databases and external APIs as needed.

[0457] Send and view high-quality output

[0458] The generated high-quality output is sent to the terminal by the server's data transmission module, and the terminal displays the received data to the user using the display module, allowing the user to accurately and quickly obtain the information they desire while protecting their privacy.

[0459] Specific examples

[0460] As a specific example, the following scenario is assumed.

[0461] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[0462] The terminal receives the input and passes it to the small-scale LLM.

[0463] 2. A small LLM preprocesses the user input.

[0464] For example, the vague expression "nearby" is made more specific, and the location information is masked to the city / ward / town / village level.

[0465] 3. The preprocessed data is sent to the server.

[0466] The data transmission module of the terminal transmits the preprocessed data to the server.

[0467] 4. The server's large-scale LLM generates high-quality restaurant lists.

[0468] The server uses a database or external API to retrieve information about nearby highly rated restaurants.

[0469] 5. The generated list is sent to the terminal and displayed to the user.

[0470] The terminal receives the list and visually presents it to the user.

[0471] In this way, the system provides high-quality, energy-efficient information while protecting user privacy.

[0472] The processing flow will be explained below.

[0473] Step 1:

[0474] A user types a question into a smartphone app: "What are some recommended restaurants nearby?"

[0475] Step 2:

[0476] The terminal receives the user's input at a data receiving module.

[0477] Step 3:

[0478] The device passes the received input to the small-scale LLM, which analyzes the question and translates it into a more specific expression. For example, it converts "nearby" to "within a 2km radius" and masks the user's current location to the city level.

[0479] Step 4:

[0480] The terminal transmits the pre-processed data to the server via a data transmission module.

[0481] Step 5:

[0482] The server receives the preprocessed data transmitted from the terminal at a data receiving module.

[0483] Step 6:

[0484] The server passes the received data to a large-scale LLM, which analyzes the user's question and generates high-quality output. For example, it may use a database or external API to collect information on highly rated restaurants within a 2km radius.

[0485] Step 7:

[0486] The server transmits high-quality output to the terminal via a data transmission module.

[0487] Step 8:

[0488] The terminal receives the high-quality output transmitted from the server at a data receiving module.

[0489] Step 9:

[0490] The data received by the terminal is visually displayed to the user in a display module, for example, a list of highly rated restaurants is displayed on the screen.

[0491] This series of processes allows users to quickly obtain high-quality information that protects their privacy in an energy-efficient manner.

[0492] Example 1

[0493] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0494] Modern information processing systems are required to quickly and accurately provide high-quality output in response to user-entered questions. However, achieving this goal poses several challenges. First, from the perspective of privacy protection, users' personal information must be handled securely. Second, if the input question is ambiguous, preprocessing is required to make it more specific. Furthermore, generating high-quality output requires processing large amounts of data, and improving energy efficiency for this purpose is also an important challenge.

[0495] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0496] In this invention, the server includes means for analyzing and preprocessing user input using a small-scale language model installed on the user terminal, means for transmitting the preprocessed data to a large-scale language model on the cloud, and means for receiving high-quality output generated by the large-scale language model, thereby enabling ambiguous questions to be specified and providing high-quality information with high energy efficiency while protecting user privacy.

[0497] A "small language model" is a language processing algorithm that runs on a terminal and analyzes and preprocesses user input.

[0498] "Preprocessing" refers to the process of removing ambiguous expressions from input data and anonymizing user privacy information as necessary.

[0499] "Anonymization" refers to making a user's information in a state where it is not possible to identify the individual in order to protect the user's privacy information.

[0500] "Concreteization" is the process of converting vague expressions into clear, specific expressions.

[0501] A "large-scale language model" is an advanced language processing algorithm that runs on the cloud and produces high-quality output.

[0502] "High-quality output" refers to relevant and accurate answers and information to users' questions.

[0503] A "database" is a system for systematically storing and managing data.

[0504] An "external API" is a program interface that connects to external databases and services to obtain information.

[0505] A "secure communication protocol" is a communication rule for safely sending and receiving data.

[0506] A "user terminal" is a computer device that can be directly operated by a user, such as a smartphone or tablet.

[0507] "Analysis" is the act of breaking down and examining input data in order to understand it and perform the necessary processing.

[0508] "Masking" is the act of hiding or making certain information invisible.

[0509] The present invention is an information processing system that combines a small-scale language model installed on a user terminal with a large-scale language model on the cloud. The system aims to provide high-quality output in response to user input. Specifically, the small-scale language model preprocesses data entered by the user, and the preprocessed data is sent to a large-scale language model on the cloud for analysis. The system then receives the high-quality output generated by the large-scale language model and displays it to the user.

[0510] Hardware and software used

[0511] User device: A computing device that can be directly operated by a user, such as a smartphone or tablet. The device is equipped with a small language model, a data reception / transmission module, and a display module.

[0512] Cloud server: A server on which a large-scale language model runs, and includes a data reception / transmission module and an information collection module.

[0513] Small language model behavior

[0514] A small language model installed on the user's device analyzes data entered by the user (e.g., "What are some recommended restaurants nearby?") and performs preprocessing. This preprocessing involves concretizing ambiguous expressions and anonymizing private information. For example, the ambiguous expression "nearby" is concretized as "within a 2km radius," and location information is converted to the city, town, or village level.

[0515] Sending data

[0516] The pre-processed data is then sent to the cloud server via the device's data transmission module, where data security is ensured using secure communication protocols such as HTTPS.

[0517] How large language models work

[0518] The preprocessed data sent to the cloud server is analyzed by a large-scale language model, which gathers information from databases and external APIs to generate the best answer to the user's question, such as a list of "highly rated restaurants within a 2km radius."

[0519] Send and view high-quality output

[0520] The generated high-quality output is sent to the user terminal by the server's data transmission module, and the terminal visually presents the received data to the user via the display module, allowing the user to obtain the information they need accurately and quickly while protecting their privacy.

[0521] Specific examples

[0522] 1. User Input: A user types into their smartphone, "What are some recommended restaurants nearby?"

[0523] 2. Preprocessing: A small language model concretizes the expression "nearby" to "within a 2km radius" and anonymizes location information to the city / ward / town / village level.

[0524] 3. Data transmission: The preprocessed data is transmitted to the cloud server using HTTPS.

[0525] 4. Parsing and output generation: A large language model uses databases and external APIs to create a list of "top rated restaurants within a 2km radius."

[0526] 5. Sending and displaying the output: The generated list is sent to the user terminal and presented to the user via the display module.

[0527] In this way, the present invention provides high-quality information while protecting the user's privacy.

[0528] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0529] Step 1: Capturing User Input

[0530] A user types "What are some recommended restaurants nearby?" into a smartphone app. The device captures this input in a data receiving module and passes it to a small language model as text data. At this stage, the input data is the raw text entered by the user.

[0531] Step 2: Preprocessing with small-scale LLM

[0532] A small language model on the device receives the captured input text and analyzes it. For the input text "What are some recommended restaurants nearby?", the ambiguous "nearby" is made more specific to "within a 2km radius," and the user's location information is further anonymized at the city / ward / town / village level. The output of this step is the preprocessed text data "What are some recommended restaurants within a 2km radius?"

[0533] Step 3: Sending preprocessed data

[0534] The preprocessed data is received by the terminal's data transmission module and sent to the cloud server. HTTPS is used for this transmission to ensure data security. The input is the preprocessed text data, and the output is the encoded data transmission.

[0535] Step 4: Receiving data on the server

[0536] The server receives data sent from the device through a data reception module, decodes the received encoded data, and prepares it for passing to a large-scale language model, where the input is the encoded preprocessed data and the output is the decoded preprocessed data.

[0537] Step 5: Generating high-quality output with large-scale LLM

[0538] The server's large-scale language model analyzes the preprocessed data. Specifically, for the question "What are some recommended restaurants within a 2km radius?", it collects information from databases and external APIs and generates a list of "highly rated restaurants within a 2km radius." The input in this step is the decoded preprocessed data, and the output is a high-quality restaurant list.

[0539] Step 6: Send high-quality output

[0540] The generated high-quality output is encoded by the server's data transmission module and sent to the terminal. Again, HTTPS is used to transmit data securely. The input is a high-quality restaurant list, and the output is the encoded restaurant list.

[0541] Step 7: Receive and display data on your device

[0542] The terminal's data receiving module receives high-quality output from the server. The received data is decoded and passed to the display module. The display module displays the information in a visually easy-to-understand format for the user, for example, as a "list of highly rated restaurants." The input in this step is the encoded restaurant list, and the output is restaurant information visually displayed to the user.

[0543] This allows users to obtain the information they need accurately and quickly while their privacy is protected.

[0544] (Application example 1)

[0545] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0546] Conventional language modeling systems have problems such as insufficient protection of user privacy information and difficulty in obtaining high-quality answers due to ambiguous user questions. Energy efficiency is also an issue. Food delivery services, in particular, require users to quickly and accurately find nearby restaurant recommendations.

[0547] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0548] In this invention, the server includes: means for analyzing and preprocessing user input using a small-scale language model; means for sending the preprocessed data to a large-scale language model on the cloud; means for receiving high-quality output generated by the large-scale language model; means for displaying the received high-quality output to the user; means for converting ambiguous expressions into specific conditions and anonymizing location information in the preprocessing; means for masking privacy information to the city / ward / town / village level to protect user privacy; means for generating a restaurant list in descending order of rating; means for collecting and organizing information based on the preprocessed data; and means for generating high-quality output from prompt sentences using a generative AI model. This enables ambiguous questions to be converted into specific, high-quality answers while protecting user privacy. It also enables users to quickly and accurately find recommended nearby restaurants for food delivery.

[0549] A "small language model" is a relatively small language processing system capable of analyzing user input and converting vague expressions into concrete terms.

[0550] "Preprocessing" is the process of analyzing user input information, concretizing ambiguous expressions, and anonymizing and masking privacy information.

[0551] A "large-scale language model in the cloud" is a large-scale language processing system accessed via the internet for complex analysis and high-quality output generation.

[0552] "High-quality output" is accurate and detailed information generated by a large-scale language model and tailored to the user's needs.

[0553] "Means for displaying to a user" refers to a display or display module for visually presenting the received output on a terminal.

[0554] "Ambiguous expressions" refer to parts of the user's input that lack specificity or are unclear.

[0555] "Specific conditions" are those that convert vague expressions into clear and detailed information.

[0556] "Means for anonymizing location information" refers to a method for protecting privacy by deleting or converting the user's current location to a level that makes it identifiable.

[0557] "Means for masking privacy information to the city, ward, town, or village level" refers to a method for preventing individuals from being identified by limiting the user's location information to the city, ward, town, or village level.

[0558] The "means for generating a restaurant list in descending order of rating" is a method for sorting nearby restaurants in a food delivery service by rating criteria.

[0559] "Means for collecting and organizing information based on preprocessed data" refers to a method for obtaining necessary information from preprocessed data and organizing it appropriately.

[0560] A "generative AI model" is an artificial intelligence language processing model that provides appropriate output based on an input prompt.

[0561] A "prompt" is an instruction or question given to an AI model to generate a specific output.

[0562] The system for implementing this invention consists of a terminal that analyzes and preprocesses user input using a small-scale language model, and a server that generates high-quality output using a large-scale language model on the cloud.

[0563] System configuration

[0564] The system consists of the following hardware and software:

[0565] Hardware

[0566] User devices (smartphones, tablets)

[0567] Cloud server with graphics card

[0568] software

[0569] Small language model (runs on device)

[0570] Large-scale language model (running on cloud servers)

[0571] Transformers Library (Hugging Face)

[0572] requests library (data communication)

[0573] React Native (user interface)

[0574] System Operation

[0575] The system operates in the following manner to efficiently process user input information and provide high quality output.

[0576] 1. Capturing User Input

[0577] A user uses a smartphone app to input a question, for example, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[0578] 2. Pretreatment by small-scale LLM

[0579] User input is first sent to a small language model on the device for preprocessing. For example, a vague expression like "nearby" is converted into a more specific condition like "within a 2km radius." Additionally, user location information is anonymized at the city / ward / town / village level, protecting user privacy.

[0580] 3. Sending preprocessed data

[0581] After the preprocessing is complete, the data is sent to the cloud server by the terminal's data transmission module, which receives the data and passes it to a large-scale language model.

[0582] 4. Generating high-quality output using large-scale LLM

[0583] A large-scale language model on a cloud server analyzes the preprocessed questions and generates high-quality output, such as a list of "top-rated restaurants within a 2km radius." The large-scale language model collects and organizes information using databases and external APIs as needed.

[0584] 5. Send and display high-quality output

[0585] The generated high-quality output is sent to the device by the cloud server's data transmission module, and the device displays the received data to the user via the display module, allowing the user to accurately and quickly obtain the information they desire while protecting their privacy.

[0586] Examples of concrete examples and prompts

[0587] A concrete example of how the system has been implemented is a food delivery application, which works as follows:

[0588] 1. A user types into the app, "What are some recommended restaurants nearby?"

[0589] 2. The app converts the input into "What are some recommended restaurants within a 2km radius?" and anonymizes the location information to "Shinjuku-ku, Tokyo" before sending it to the server.

[0590] 3. The server uses an AI model to generate a list of highly rated restaurants nearby.

[0591] 4. Send the list to the app and display it to the user.

[0592] Example prompt sentence:

[0593] Please list recommended restaurants within a 2km radius. The user's anonymized location is "Shinjuku-ku, Tokyo."

[0594] In this way, it is possible to provide highly accurate restaurant information for food delivery while ensuring the user's privacy.

[0595] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0596] Step 1:

[0597] A user opens a smartphone application and inputs the question, "What are some recommended restaurants nearby?" This input is captured on the user's device and passed to a data receiving module. The input data is natural language text, and the operation to capture the input data is performed via a user interface. In this case, the input is "What are some recommended restaurants nearby?" and the output is the user's input text data.

[0598] Step 2:

[0599] The device passes the captured user input to a small-scale language model and begins preprocessing. Specifically, the small-scale LLM concretizes the ambiguous expression "nearby" to "within a 2km radius" and anonymizes the user's location information to "Shinjuku Ward, Tokyo" at the city / ward / town / village level. The input for preprocessing is the captured user's natural language text and location information, and the output is a preprocessed concretized question: "Recommend restaurants within a 2km radius (Shinjuku Ward, Tokyo)."

[0600] Step 3:

[0601] The preprocessed data is sent to the cloud server by the terminal's data transmission module. During the transmission process, the data is encrypted using a protocol such as HTTPS request and sent to the server. The input is the preprocessed query data, and the output is the completion of data transmission to the cloud server.

[0602] Step 4:

[0603] The cloud server passes the received preprocessed data to a large-scale language model (generative AI model). The large-scale LLM uses external databases and APIs to provide high-quality information. Specifically, it generates a list of "highly rated restaurants within a 2km radius." The input of the large-scale LLM is the preprocessed, specified question data, and the output is a list of highly rated restaurants.

[0604] Step 5:

[0605] The server receives the high-quality output (restaurant list) generated by the large-scale LLM and sends it to the terminal. The transmission process is also encrypted and is carried out using protocols such as HTTPS requests. The input here is the generated restaurant list data, and the output is the completion of transmission of the restaurant list to the terminal.

[0606] Step 6:

[0607] The terminal passes the restaurant list received from the server to the display module, which then visually presents it to the user. Specifically, the restaurant list is displayed on the application interface using a framework such as React Native. The input here is the received restaurant list data, and the output is the specific restaurant list visually displayed to the user.

[0608] This series of processing steps allows users to quickly obtain high-quality restaurant information for food delivery services while protecting their privacy.

[0609] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0610] This invention is a system that combines a small-scale language model (hereinafter referred to as small-scale LLM) with a large-scale language model (hereinafter referred to as large-scale LLM), and adds an emotion engine that recognizes user emotions. It efficiently processes user input and provides high-quality output according to the user's emotions. The system aims to protect privacy, improve energy efficiency, and respond to user emotions.

[0611] System configuration

[0612] The system is accessed by users through specific input devices (such as smartphones or tablets) and includes the following components:

[0613] Terminal

[0614] Small scale LLM

[0615] Emotion Engine

[0616] Data reception and transmission module

[0617] Display Module

[0618] server

[0619] large scale LLM

[0620] Data reception and transmission module

[0621] Information Collection Module

[0622] System Operation

[0623] When a user uses the system to ask a question, the following process is carried out.

[0624] Capturing User Input

[0625] A user inputs a question into a smartphone app. For example, they input a question like, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[0626] Emotion recognition with emotion engine

[0627] The device sends the captured input to an emotion engine to recognize the user's emotions, for example, determining whether the user is in a hurry or relaxed state based on the style and tone of the question.

[0628] Pretreatment by small-scale LLM

[0629] The device uses the emotion information obtained from the emotion engine and passes it to the small-scale LLM. The small-scale LLM analyzes the question and converts it into a more specific expression. For example, it converts the vague expression "nearby" into "within a 2km radius" and masks the user's current location information to the city / ward / town / village level. If the device detects a sense of urgency, it sets a priority to speed up processing.

[0630] Sending preprocessed data

[0631] The pre-processed data is sent to the server by the terminal's data transmission module, which receives the data and passes it to the large-scale LLM.

[0632] Generating high-quality output with large-scale LLM

[0633] The server-side large-scale LLM analyzes the preprocessed queries and generates high-quality output that meets the user's requirements. For example, it generates a list of "highly rated restaurants within a 2km radius." The large-scale LLM collects and organizes information using databases and external APIs as needed.

[0634] Send and view high-quality output

[0635] The generated high-quality output is transmitted to the terminal by the server's data transmission module, and the terminal displays the received data to the user using the display module. This allows the user to quickly obtain high-quality information according to their emotions while protecting their privacy.

[0636] Specific examples

[0637] As a specific example, the following scenario is assumed.

[0638] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[0639] The device receives the input and passes it to the emotion engine.

[0640] 2. The emotion engine recognizes the user's emotions.

[0641] For example, it is determined that the user is in a hurry.

[0642] 3. A small-scale LLM performs preprocessing based on emotion information.

[0643] It clarifies the vague term "nearby" by masking location information to city level, and sets priorities for faster processing based on feelings of urgency.

[0644] 4. The preprocessed data is sent to the server.

[0645] The data transmission module of the terminal transmits the preprocessed data to the server.

[0646] 5. The server's large-scale LLM generates high-quality restaurant lists.

[0647] The server uses a database or external API to collect information about nearby highly rated restaurants.

[0648] 6. The generated list is sent to the terminal and displayed to the user.

[0649] The terminal receives the list and visually presents it to the user.

[0650] In this way, the system understands users' emotions while providing high-quality information with improved energy efficiency and privacy protection.

[0651] The processing flow will be explained below.

[0652] Step 1:

[0653] A user types a question into a smartphone app: "What are some recommended restaurants nearby?"

[0654] Step 2:

[0655] The terminal receives the user's input at a data receiving module.

[0656] Step 3:

[0657] The device sends the received input to an emotion engine that analyzes the user's emotions, such as identifying emotions like "hurried" or "relaxed" based on the user's writing style and tone.

[0658] Step 4:

[0659] If the device receives the analysis results from the emotion engine and determines that the user is in a hurry, it sends the input data to a small-scale LLM and instructs it to perform high-priority processing.

[0660] Step 5:

[0661] A small LLM analyzes user input and converts vague expressions into specific ones, for example, changing "nearby" to "within a 2km radius" and masking the user's current location to city level.

[0662] Step 6:

[0663] The terminal transmits the pre-processed data to the server via a data transmission module.

[0664] Step 7:

[0665] The server receives the preprocessed data transmitted from the terminal at a data receiving module.

[0666] Step 8:

[0667] The server passes the received data to a large-scale LLM, which analyzes the data and generates high-quality output based on the user's sentiment and requirements. For example, it creates a list of highly rated restaurants within a 2km radius.

[0668] Step 9:

[0669] The high-quality output generated by the server is transmitted to the terminal by a data transmission module.

[0670] Step 10:

[0671] The terminal receives the high-quality output transmitted from the server at a data receiving module.

[0672] Step 11:

[0673] The data received by the terminal is visually displayed to the user in a display module, for example, a list of highly rated restaurants is displayed on the screen.

[0674] Through this series of steps, users can quickly obtain high-quality information that corresponds to their emotions while their privacy is protected.

[0675] Example 2

[0676] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0677] Current language model systems perform uniform processing without considering the user's emotions, resulting in a poor user experience. They also sometimes fail to adequately protect the user's privacy and can be slow to respond. Furthermore, they lack the ability to properly concretize ambiguous input, making it difficult to provide the accurate information the user desires.

[0678] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for capturing user input, means for analyzing and preprocessing the user input using a small-scale language model, means for performing sentiment analysis, means for transmitting the preprocessed data to a large-scale language model on the cloud, means for receiving high-quality output generated by the large-scale language model, means for displaying the received high-quality output to the user, and means for concretizing ambiguous expressions based on the user input and setting priorities if necessary. This enables the provision of high-quality information that takes user emotions into consideration, protection of privacy, and highly efficient data processing.

[0679] A "small language model" is a computer program that analyzes and preprocesses user input data, optimized to operate in environments with limited computing resources.

[0680] A "large-scale language model" is a computer program that uses large amounts of data and computing resources to perform advanced natural language processing and generate high-quality output.

[0681] "Preprocessing" refers to a series of operations that analyze data entered by a user and perform masking to concretize ambiguous expressions and protect privacy.

[0682] "Receiving" means receiving data sent from a sender.

[0683] "Emotion analysis" is the process of inferring a user's emotions and intentions based on the data they input.

[0684] "Data transmission" means sending pre-processed data to a designated destination.

[0685] "High-quality output" is information generated by a large-scale language model to accurately respond to user requests.

[0686] A "data receiving module" is hardware or software that has the function of receiving data from the outside and passing it on to internal processing.

[0687] A "data transmission module" is hardware or software that has the function of transmitting pre-processed or generated data to the outside.

[0688] A "display module" is hardware or software that has the function of visually displaying received information to a user.

[0689] This system combines small-scale and large-scale language models and adds an emotion analysis engine to recognize user emotions. It efficiently processes user input and provides high-quality output that reflects the user's emotions. This system aims to protect privacy, improve energy efficiency, and respond to user emotions.

[0690] System configuration

[0691] The system is accessed by users through specific input devices (such as smartphones or tablets) and includes the following components:

[0692] Terminal

[0693] Small language models

[0694] It is a computer program that analyzes and preprocesses user input data.

[0695] Sentiment Analysis Engine

[0696] It is a computer program that analyzes emotions and intentions based on user input data.

[0697] Data Receiving Module

[0698] It has the function of receiving input from the user and passing it on to internal processing.

[0699] Data Transmission Module

[0700] It has a function for sending preprocessed data to the server.

[0701] Display Module

[0702] It has the ability to visually display high quality output received from the server to the user.

[0703] server

[0704] Large-scale language models

[0705] It is a computer program that performs advanced natural language processing and generates high-quality output that meets user requirements.

[0706] Data Receiving Module

[0707] It has the function of receiving data sent from the terminal and passing it on to internal processing.

[0708] Information Collection Module

[0709] It has the ability to collect information from databases and external APIs as needed and organize the data.

[0710] Specific examples

[0711] When a user asks a question

[0712] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[0713] The terminal's data receiving module captures this input.

[0714] 2. The sentiment analysis engine recognizes the user's emotions.

[0715] For example, the style and tone of the question can determine whether the user is in a hurry or relaxed.

[0716] 3. A small language model performs preprocessing based on emotion information.

[0717] The vague expression "nearby" is specified as "within a 2km radius," location information is masked to the city level, and priority is set for fast processing based on the feeling of urgency.

[0718] 4. The preprocessed data is sent to the server.

[0719] The data transmission module of the terminal transmits the preprocessed data to the server.

[0720] 5. The server's large-scale language model generates high-quality restaurant lists.

[0721] A large language model uses databases and external APIs to create a list of "highly rated restaurants within a 2km radius."

[0722] 6. The generated list is sent to the terminal and displayed to the user.

[0723] The terminal receives the list and visually presents it to the user using a display module.

[0724] Examples of prompt statements

[0725] For example, if a user inputs a question such as "What are some recommended restaurants nearby?", the above process will promptly provide high-quality information while taking into consideration the user's feelings.

[0726] In this way, this system understands the user's emotions and provides high-quality information while protecting privacy.

[0727] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0728] Step 1: Capturing User Input

[0729] A user types into their smartphone, "What are some recommended restaurants near me?"

[0730] The terminal's data reception module captures this input. Specifically, it acquires the string entered in the text box as internal data.

[0731] Input: User question (e.g., "What are some recommended restaurants nearby?")

[0732] Output: Captured user question data (e.g., "What are some recommended restaurants nearby?")

[0733] Step 2: Emotion recognition by the emotion engine

[0734] The device sends the captured user question to the emotion engine, which analyzes emotions from the input text.

[0735] Input: User question data (e.g., "What are some recommended restaurants nearby?")

[0736] Output: Emotional information (e.g., user is in a hurry)

[0737] Step 3: Preprocessing with small-scale LLM

[0738] The device passes the captured question data based on the emotion information to the small-scale LLM, which analyzes the question and makes ambiguous expressions concrete.

[0739] Input: User question data and emotion information (e.g., "What are some recommended restaurants nearby?", emotion: "I'm in a hurry")

[0740] Specific behavior: Converts the vague expression "nearby" to "within a 2km radius," masks location information to the city / ward / town / village level for privacy reasons, and sets priorities for faster processing.

[0741] Output: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", sentiment: in a hurry)

[0742] Step 4: Sending preprocessed data

[0743] The data transmission module of the terminal transmits the preprocessed data to the server, specifically, by using an HTTP request.

[0744] Input: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", Sentiment: Hurry)

[0745] Output: Data sent to the server

[0746] Step 5: Generating high-quality output with large-scale LLM

[0747] The server receives the pre-processed data and passes it to the large-scale LLM, which uses databases and external APIs to generate high-quality output.

[0748] Input: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", Sentiment: Hurry)

[0749] Specific behavior: Performs database queries and external API calls, and generates a list of restaurants based on the obtained data.

[0750] Output: High-quality output data (e.g., "List of highly rated restaurants within a 2km radius")

[0751] Step 6: Send and view high-quality output

[0752] The data transmission module of the server transmits the generated high-quality output to the terminal.

[0753] The terminal receives the output sent from the server and visually displays it to the user using a display module.

[0754] Input: High-quality output data (e.g., "List of highly rated restaurants within a 2km radius")

[0755] Specific behavior: Renders the received data in the UI and displays it to the user.

[0756] Output: High-quality information visually presented to the user

[0757] In this way, the system understands users' emotions while providing high-quality information with improved energy efficiency and privacy protection.

[0758] (Application example 2)

[0759] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0760] Conventional systems have had difficulty in accurately recognizing emotions in response to user input and providing personalized, high-quality information based on that emotion. Furthermore, while there is a need to efficiently provide information tailored to the user's emotions, there are insufficient measures to protect privacy and prevent unnecessary repetition of questions. Understanding customer emotions and providing real-time services based on those emotions is particularly important in brick-and-mortar stores, but conventional technologies have had difficulty meeting this requirement.

[0761] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing and preprocessing a user's input using a small-scale language model, means for transmitting the preprocessed data to a large-scale language model on the cloud, means for receiving high-quality output generated by the large-scale language model, means for displaying the received high-quality output to the user, means for recognizing emotions from the user's input, and means for adjusting data preprocessing based on the recognized emotions. This makes it possible to provide information according to the user's emotions in real time, protect privacy, and prevent unnecessary repetition of questions.

[0762] A "small language model" is a lightweight, energy-efficient natural language processing model used to parse and preprocess user input.

[0763] A "large-scale language model" is a high-performance natural language processing model that runs on the cloud and generates high-quality output based on pre-processed data.

[0764] "User input" refers to data such as text information and voice information that a user provides to the system.

[0765] "Preprocessing" refers to the initial processing of user input to parse it and convert it into a format that meets specific needs.

[0766] "Sending to the cloud" refers to transferring data from the user's terminal to a remote server.

[0767] "High-quality output" refers to highly accurate information and results generated by a large-scale language model that are tailored to the user's needs.

[0768] "Receiving" refers to receiving the generated output from the server at the user's terminal.

[0769] "Displaying" refers to visually presenting the received high quality output to a user.

[0770] "Emotion recognition" refers to the process of determining a user's emotional state from their input.

[0771] "Adjusting data preprocessing" refers to changing and optimizing the content and methods of preprocessing based on the recognized emotions.

[0772] The system for implementing this invention analyzes user input, recognizes emotions, and provides high-quality output based on those emotions. The system consists of two main components: a terminal and a server.

[0773] Program Overview

[0774] 1. Capturing user input

[0775] The user inputs data through a terminal (e.g., a smartphone). The input data can be text data or voice data.

[0776] 2. Emotional Recognition

[0777] The device passes the input data to an emotion engine (e.g., Hugging Face sentiment-analysis) to recognize the user's emotion. For example, if the user input is "Are there any deals?", the device recognizes the user's emotion as "happy" based on that input.

[0778] 3. Pretreatment

[0779] It uses a small language model (e.g., T5-small) to analyze user input and convert it into a concrete form while taking into account emotional information, for example, converting the vague expression "great deal" into "today's sale items," and adjusts the processing speed based on the emotion.

[0780] 4. Data transmission

[0781] The pre-processed data is transmitted to the server by the data transmission module of the terminal.

[0782] 5. Processing with large-scale language models

[0783] The server uses large-scale language models to analyze the pre-processed data and generate high-quality output, such as a list of today's sale items.

[0784] 6. Send and display high-quality output

[0785] The server generates high-quality output and sends it to the terminal, which visually displays the received data to the user.

[0786] Hardware and software used

[0787] Device: User device such as a smartphone or tablet

[0788] Emotion Engine: A sentiment-analysis model of Hugging Face

[0789] Small language models: Natural language processing models such as T5-small

[0790] Server: High-performance server on the cloud

[0791] Large-scale language model: High-performance natural language processing model running on a server

[0792] Specific examples

[0793] The specific scenario is as follows:

[0794] 1. The user speaks into the smartphone app and says, "Do you have any deals?"

[0795] 2. The device captures the input and recognizes the "happy" emotion using its emotion engine.

[0796] 3. Small LLMs concretize the expression "good deal" and transform it into "What's on sale today?"

[0797] 4. Send it to the server, and the large-scale LLM generates a "list of today's sale items" that meets your request.

[0798] 5. Once the list is returned, the device displays it to the user along with "Today's Sale Items."

[0799] Example prompt sentence:

[0800] Q: Are there any great deals? I'd be happy to provide information at a speed that matches my needs. A:

[0801] In this way, the system can provide personalized information based on user sentiment in real time, improving the customer experience.

[0802] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0803] Step 1:

[0804] The user types "Do you have any deals?" into the smartphone app. The input can be text or voice. The device captures this data and stores it as input data.

[0805] Step 2:

[0806] The device passes the captured input data to the emotion engine, which (Hugging Face's sentiment-analysis model) recognizes the emotion from the user's input. In this step, based on the input data (e.g., "Are there any deals available?"), the emotion engine outputs the emotion information "happy."

[0807] Step 3:

[0808] The device takes the recognized emotional information into consideration and passes the user's input to a small-scale LLM (T5-small model). The small-scale LLM analyzes the input and performs data preprocessing based on the emotional information. In this step, the device outputs preprocessed data (e.g., "What are the sale items today?") based on the input data and emotional information (e.g., "Are there any deals?" and "I'm happy").

[0809] Step 4:

[0810] The data transmission module of the device sends the preprocessed data to a server on the cloud, where the preprocessed data ("What are the sale items today?") is transferred to the server as input data.

[0811] Step 5:

[0812] The server passes the received preprocessed data to the large-scale LLM, which analyzes the preprocessed data and generates the corresponding high-quality output. In this step, the server outputs a high-quality output ("List of sale items") based on the preprocessed input data ("What items are on sale today?").

[0813] Step 6:

[0814] The server's data transmission module transmits the generated high-quality output to the terminal, where the generated data ("list of sale items") is transmitted to the terminal.

[0815] Step 7:

[0816] The terminal uses a display module to display the received high-quality data to the user. In this step, the terminal visually presents the high-quality output data ("list of sale items") to the user.

[0817] The above is the specific process flow from user input to the provision of high-quality information. Data is input and output at each step, and the next step is executed based on the results, thereby realizing the provision of highly accurate, emotion-responsive information.

[0818] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0819] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0820] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0821] [Third embodiment]

[0822] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0823] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0824] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0825] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0826] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0827] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0828] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0829] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0830] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0831] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0832] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0833] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0834] The present invention is a system that combines small-scale language models (hereinafter referred to as small-scale LLMs) and large-scale language models (hereinafter referred to as large-scale LLMs), which enables efficient processing of user input and provides high-quality output. This system aims to protect privacy and improve energy efficiency, among other things.

[0835] System configuration

[0836] This system is intended to be used by users through devices such as smartphones and tablets. The system includes the following components:

[0837] Terminal

[0838] Small scale LLM

[0839] Data reception and transmission module

[0840] Display Module

[0841] server

[0842] large scale LLM

[0843] Data reception and transmission module

[0844] Information Collection Module

[0845] System Operation

[0846] When a user uses the system to ask a question, the following process is carried out.

[0847] Capturing User Input

[0848] A user inputs a question into a smartphone app. For example, they input a question like, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[0849] Pretreatment by small-scale LLM

[0850] The device first sends the captured input to a small-scale LLM for preprocessing. Specifically, it removes ambiguous expressions from the question and masks the user's privacy information (e.g., current location information). For example, it clarifies the expression "nearby" to "within a 2km radius." This preprocessing aims to clarify the question while protecting the user's privacy.

[0851] Sending preprocessed data

[0852] The pre-processed data is sent to the server by the terminal's data transmission module, which receives the data and passes it to the large-scale LLM.

[0853] Generating high-quality output with large-scale LLM

[0854] The server-side large-scale LLM analyzes the preprocessed queries and generates high-quality output that meets the user's requirements. For example, it generates a list of "highly rated restaurants within a 2km radius." The large-scale LLM collects and organizes information using databases and external APIs as needed.

[0855] Send and view high-quality output

[0856] The generated high-quality output is sent to the terminal by the server's data transmission module, and the terminal displays the received data to the user using the display module, allowing the user to accurately and quickly obtain the information they desire while protecting their privacy.

[0857] Specific examples

[0858] As a specific example, the following scenario is assumed.

[0859] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[0860] The terminal receives the input and passes it to the small-scale LLM.

[0861] 2. A small LLM preprocesses the user input.

[0862] For example, the vague expression "nearby" is made more specific, and the location information is masked to the city / ward / town / village level.

[0863] 3. The preprocessed data is sent to the server.

[0864] The data transmission module of the terminal transmits the preprocessed data to the server.

[0865] 4. The server's large-scale LLM generates high-quality restaurant lists.

[0866] The server uses a database or external API to retrieve information about nearby highly rated restaurants.

[0867] 5. The generated list is sent to the terminal and displayed to the user.

[0868] The terminal receives the list and visually presents it to the user.

[0869] In this way, the system provides high-quality, energy-efficient information while protecting user privacy.

[0870] The processing flow will be explained below.

[0871] Step 1:

[0872] A user types a question into a smartphone app: "What are some recommended restaurants nearby?"

[0873] Step 2:

[0874] The terminal receives the user's input at a data receiving module.

[0875] Step 3:

[0876] The device passes the received input to the small-scale LLM, which analyzes the question and translates it into a more specific expression. For example, it converts "nearby" to "within a 2km radius" and masks the user's current location to the city level.

[0877] Step 4:

[0878] The terminal transmits the pre-processed data to the server via a data transmission module.

[0879] Step 5:

[0880] The server receives the preprocessed data transmitted from the terminal at a data receiving module.

[0881] Step 6:

[0882] The server passes the received data to a large-scale LLM, which analyzes the user's question and generates high-quality output. For example, it may use a database or external API to collect information on highly rated restaurants within a 2km radius.

[0883] Step 7:

[0884] The server transmits high-quality output to the terminal via a data transmission module.

[0885] Step 8:

[0886] The terminal receives the high-quality output transmitted from the server at a data receiving module.

[0887] Step 9:

[0888] The data received by the terminal is visually displayed to the user in a display module, for example, a list of highly rated restaurants is displayed on the screen.

[0889] This series of processes allows users to quickly obtain high-quality information that protects their privacy in an energy-efficient manner.

[0890] Example 1

[0891] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0892] Modern information processing systems are required to quickly and accurately provide high-quality output in response to user-entered questions. However, achieving this goal poses several challenges. First, from the perspective of privacy protection, users' personal information must be handled securely. Second, if the input question is ambiguous, preprocessing is required to make it more specific. Furthermore, generating high-quality output requires processing large amounts of data, and improving energy efficiency for this purpose is also an important challenge.

[0893] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0894] In this invention, the server includes means for analyzing and preprocessing user input using a small-scale language model installed on the user terminal, means for transmitting the preprocessed data to a large-scale language model on the cloud, and means for receiving high-quality output generated by the large-scale language model, thereby enabling ambiguous questions to be specified and providing high-quality information with high energy efficiency while protecting user privacy.

[0895] A "small language model" is a language processing algorithm that runs on a terminal and analyzes and preprocesses user input.

[0896] "Preprocessing" refers to the process of removing ambiguous expressions from input data and anonymizing user privacy information as necessary.

[0897] "Anonymization" refers to making a user's information in a state where it is not possible to identify the individual in order to protect the user's privacy information.

[0898] "Concreteization" is the process of converting vague expressions into clear, specific expressions.

[0899] A "large-scale language model" is an advanced language processing algorithm that runs on the cloud and produces high-quality output.

[0900] "High-quality output" refers to relevant and accurate answers and information to users' questions.

[0901] A "database" is a system for systematically storing and managing data.

[0902] An "external API" is a program interface that connects to external databases and services to obtain information.

[0903] A "secure communication protocol" is a communication rule for safely sending and receiving data.

[0904] A "user terminal" is a computer device that can be directly operated by a user, such as a smartphone or tablet.

[0905] "Analysis" is the act of breaking down and examining input data in order to understand it and perform the necessary processing.

[0906] "Masking" is the act of hiding or making certain information invisible.

[0907] The present invention is an information processing system that combines a small-scale language model installed on a user terminal with a large-scale language model on the cloud. The system aims to provide high-quality output in response to user input. Specifically, the small-scale language model preprocesses data entered by the user, and the preprocessed data is sent to a large-scale language model on the cloud for analysis. The system then receives the high-quality output generated by the large-scale language model and displays it to the user.

[0908] Hardware and software used

[0909] User device: A computing device that can be directly operated by a user, such as a smartphone or tablet. The device is equipped with a small language model, a data reception / transmission module, and a display module.

[0910] Cloud server: A server on which a large-scale language model runs, and includes a data reception / transmission module and an information collection module.

[0911] Small language model behavior

[0912] A small language model installed on the user's device analyzes data entered by the user (e.g., "What are some recommended restaurants nearby?") and performs preprocessing. This preprocessing involves concretizing ambiguous expressions and anonymizing private information. For example, the ambiguous expression "nearby" is concretized as "within a 2km radius," and location information is converted to the city, town, or village level.

[0913] Sending data

[0914] The pre-processed data is then sent to the cloud server via the device's data transmission module, where data security is ensured using secure communication protocols such as HTTPS.

[0915] How large language models work

[0916] The preprocessed data sent to the cloud server is analyzed by a large-scale language model, which gathers information from databases and external APIs to generate the best answer to the user's question, such as a list of "highly rated restaurants within a 2km radius."

[0917] Send and view high-quality output

[0918] The generated high-quality output is sent to the user terminal by the server's data transmission module, and the terminal visually presents the received data to the user via the display module, allowing the user to obtain the information they need accurately and quickly while protecting their privacy.

[0919] Specific examples

[0920] 1. User Input: A user types into their smartphone, "What are some recommended restaurants nearby?"

[0921] 2. Preprocessing: A small language model concretizes the expression "nearby" to "within a 2km radius" and anonymizes location information to the city / ward / town / village level.

[0922] 3. Data transmission: The preprocessed data is transmitted to the cloud server using HTTPS.

[0923] 4. Parsing and output generation: A large language model uses databases and external APIs to create a list of "top rated restaurants within a 2km radius."

[0924] 5. Sending and displaying the output: The generated list is sent to the user terminal and presented to the user via the display module.

[0925] In this way, the present invention provides high-quality information while protecting the user's privacy.

[0926] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0927] Step 1: Capturing User Input

[0928] A user types "What are some recommended restaurants nearby?" into a smartphone app. The device captures this input in a data receiving module and passes it to a small language model as text data. At this stage, the input data is the raw text entered by the user.

[0929] Step 2: Preprocessing with small-scale LLM

[0930] A small language model on the device receives the captured input text and analyzes it. For the input text "What are some recommended restaurants nearby?", the ambiguous "nearby" is made more specific to "within a 2km radius," and the user's location information is further anonymized at the city / ward / town / village level. The output of this step is the preprocessed text data "What are some recommended restaurants within a 2km radius?"

[0931] Step 3: Sending preprocessed data

[0932] The preprocessed data is received by the terminal's data transmission module and sent to the cloud server. HTTPS is used for this transmission to ensure data security. The input is the preprocessed text data, and the output is the encoded data transmission.

[0933] Step 4: Receiving data on the server

[0934] The server receives data sent from the device through a data reception module, decodes the received encoded data, and prepares it for passing to a large-scale language model, where the input is the encoded preprocessed data and the output is the decoded preprocessed data.

[0935] Step 5: Generating high-quality output with large-scale LLM

[0936] The server's large-scale language model analyzes the preprocessed data. Specifically, for the question "What are some recommended restaurants within a 2km radius?", it collects information from databases and external APIs and generates a list of "highly rated restaurants within a 2km radius." The input in this step is the decoded preprocessed data, and the output is a high-quality restaurant list.

[0937] Step 6: Send high-quality output

[0938] The generated high-quality output is encoded by the server's data transmission module and sent to the terminal. Again, HTTPS is used to transmit data securely. The input is a high-quality restaurant list, and the output is the encoded restaurant list.

[0939] Step 7: Receive and display data on your device

[0940] The terminal's data receiving module receives high-quality output from the server. The received data is decoded and passed to the display module. The display module displays the information in a visually easy-to-understand format for the user, for example, as a "list of highly rated restaurants." The input in this step is the encoded restaurant list, and the output is restaurant information visually displayed to the user.

[0941] This allows users to obtain the information they need accurately and quickly while their privacy is protected.

[0942] (Application example 1)

[0943] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0944] Conventional language modeling systems have problems such as insufficient protection of user privacy information and difficulty in obtaining high-quality answers due to ambiguous user questions. Energy efficiency is also an issue. Food delivery services, in particular, require users to quickly and accurately find nearby restaurant recommendations.

[0945] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0946] In this invention, the server includes: means for analyzing and preprocessing user input using a small-scale language model; means for sending the preprocessed data to a large-scale language model on the cloud; means for receiving high-quality output generated by the large-scale language model; means for displaying the received high-quality output to the user; means for converting ambiguous expressions into specific conditions and anonymizing location information in the preprocessing; means for masking privacy information to the city / ward / town / village level to protect user privacy; means for generating a restaurant list in descending order of rating; means for collecting and organizing information based on the preprocessed data; and means for generating high-quality output from prompt sentences using a generative AI model. This enables ambiguous questions to be converted into specific, high-quality answers while protecting user privacy. It also enables users to quickly and accurately find recommended nearby restaurants for food delivery.

[0947] A "small language model" is a relatively small language processing system capable of analyzing user input and converting vague expressions into concrete terms.

[0948] "Preprocessing" is the process of analyzing user input information, concretizing ambiguous expressions, and anonymizing and masking privacy information.

[0949] A "large-scale language model in the cloud" is a large-scale language processing system accessed via the internet for complex analysis and high-quality output generation.

[0950] "High-quality output" is accurate and detailed information generated by a large-scale language model and tailored to the user's needs.

[0951] "Means for displaying to a user" refers to a display or display module for visually presenting the received output on a terminal.

[0952] "Ambiguous expressions" refer to parts of the user's input that lack specificity or are unclear.

[0953] "Specific conditions" are those that convert vague expressions into clear and detailed information.

[0954] "Means for anonymizing location information" refers to a method for protecting privacy by deleting or converting the user's current location to a level that makes it identifiable.

[0955] "Means for masking privacy information to the city, ward, town, or village level" refers to a method for preventing individuals from being identified by limiting the user's location information to the city, ward, town, or village level.

[0956] The "means for generating a restaurant list in descending order of rating" is a method for sorting nearby restaurants in a food delivery service by rating criteria.

[0957] "Means for collecting and organizing information based on preprocessed data" refers to a method for obtaining necessary information from preprocessed data and organizing it appropriately.

[0958] A "generative AI model" is an artificial intelligence language processing model that provides appropriate output based on an input prompt.

[0959] A "prompt" is an instruction or question given to an AI model to generate a specific output.

[0960] The system for implementing this invention consists of a terminal that analyzes and preprocesses user input using a small-scale language model, and a server that generates high-quality output using a large-scale language model on the cloud.

[0961] System configuration

[0962] The system consists of the following hardware and software:

[0963] Hardware

[0964] User devices (smartphones, tablets)

[0965] Cloud server with graphics card

[0966] software

[0967] Small language model (runs on device)

[0968] Large-scale language model (running on cloud servers)

[0969] Transformers Library (Hugging Face)

[0970] requests library (data communication)

[0971] React Native (user interface)

[0972] System Operation

[0973] The system operates in the following manner to efficiently process user input information and provide high quality output.

[0974] 1. Capturing User Input

[0975] A user uses a smartphone app to input a question, for example, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[0976] 2. Pretreatment by small-scale LLM

[0977] User input is first sent to a small language model on the device for preprocessing. For example, a vague expression like "nearby" is converted into a more specific condition like "within a 2km radius." Additionally, user location information is anonymized at the city / ward / town / village level, protecting user privacy.

[0978] 3. Sending preprocessed data

[0979] After the preprocessing is complete, the data is sent to the cloud server by the terminal's data transmission module, which receives the data and passes it to a large-scale language model.

[0980] 4. Generating high-quality output using large-scale LLM

[0981] A large-scale language model on a cloud server analyzes the preprocessed questions and generates high-quality output, such as a list of "top-rated restaurants within a 2km radius." The large-scale language model collects and organizes information using databases and external APIs as needed.

[0982] 5. Send and display high-quality output

[0983] The generated high-quality output is sent to the device by the cloud server's data transmission module, and the device displays the received data to the user via the display module, allowing the user to accurately and quickly obtain the information they desire while protecting their privacy.

[0984] Examples of concrete examples and prompts

[0985] A concrete example of how the system has been implemented is a food delivery application, which works as follows:

[0986] 1. A user types into the app, "What are some recommended restaurants nearby?"

[0987] 2. The app converts the input into "What are some recommended restaurants within a 2km radius?" and anonymizes the location information to "Shinjuku-ku, Tokyo" before sending it to the server.

[0988] 3. The server uses an AI model to generate a list of highly rated restaurants nearby.

[0989] 4. Send the list to the app and display it to the user.

[0990] Example prompt sentence:

[0991] Please list recommended restaurants within a 2km radius. The user's anonymized location is "Shinjuku-ku, Tokyo."

[0992] In this way, it is possible to provide highly accurate restaurant information for food delivery while ensuring the user's privacy.

[0993] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0994] Step 1:

[0995] A user opens a smartphone application and inputs the question, "What are some recommended restaurants nearby?" This input is captured on the user's device and passed to a data receiving module. The input data is natural language text, and the operation to capture the input data is performed via a user interface. In this case, the input is "What are some recommended restaurants nearby?" and the output is the user's input text data.

[0996] Step 2:

[0997] The device passes the captured user input to a small-scale language model and begins preprocessing. Specifically, the small-scale LLM concretizes the ambiguous expression "nearby" to "within a 2km radius" and anonymizes the user's location information to "Shinjuku Ward, Tokyo" at the city / ward / town / village level. The input for preprocessing is the captured user's natural language text and location information, and the output is a preprocessed concretized question: "Recommend restaurants within a 2km radius (Shinjuku Ward, Tokyo)."

[0998] Step 3:

[0999] The preprocessed data is sent to the cloud server by the terminal's data transmission module. During the transmission process, the data is encrypted using a protocol such as HTTPS request and sent to the server. The input is the preprocessed query data, and the output is the completion of data transmission to the cloud server.

[1000] Step 4:

[1001] The cloud server passes the received preprocessed data to a large-scale language model (generative AI model). The large-scale LLM uses external databases and APIs to provide high-quality information. Specifically, it generates a list of "highly rated restaurants within a 2km radius." The input of the large-scale LLM is the preprocessed, specified question data, and the output is a list of highly rated restaurants.

[1002] Step 5:

[1003] The server receives the high-quality output (restaurant list) generated by the large-scale LLM and sends it to the terminal. The transmission process is also encrypted and is carried out using protocols such as HTTPS requests. The input here is the generated restaurant list data, and the output is the completion of transmission of the restaurant list to the terminal.

[1004] Step 6:

[1005] The terminal passes the restaurant list received from the server to the display module, which then visually presents it to the user. Specifically, the restaurant list is displayed on the application interface using a framework such as React Native. The input here is the received restaurant list data, and the output is the specific restaurant list visually displayed to the user.

[1006] This series of processing steps allows users to quickly obtain high-quality restaurant information for food delivery services while protecting their privacy.

[1007] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1008] This invention is a system that combines a small-scale language model (hereinafter referred to as small-scale LLM) with a large-scale language model (hereinafter referred to as large-scale LLM), and adds an emotion engine that recognizes user emotions. It efficiently processes user input and provides high-quality output according to the user's emotions. The system aims to protect privacy, improve energy efficiency, and respond to user emotions.

[1009] System configuration

[1010] The system is accessed by users through specific input devices (such as smartphones or tablets) and includes the following components:

[1011] Terminal

[1012] Small scale LLM

[1013] Emotion Engine

[1014] Data reception and transmission module

[1015] Display Module

[1016] server

[1017] large scale LLM

[1018] Data reception and transmission module

[1019] Information Collection Module

[1020] System Operation

[1021] When a user uses the system to ask a question, the following process is carried out.

[1022] Capturing User Input

[1023] A user inputs a question into a smartphone app. For example, they input a question like, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[1024] Emotion recognition with emotion engine

[1025] The device sends the captured input to an emotion engine to recognize the user's emotions, for example, determining whether the user is in a hurry or relaxed state based on the style and tone of the question.

[1026] Pretreatment by small-scale LLM

[1027] The device uses the emotion information obtained from the emotion engine and passes it to the small-scale LLM. The small-scale LLM analyzes the question and converts it into a more specific expression. For example, it converts the vague expression "nearby" into "within a 2km radius" and masks the user's current location information to the city / ward / town / village level. If the device detects a sense of urgency, it sets a priority to speed up processing.

[1028] Sending preprocessed data

[1029] The pre-processed data is sent to the server by the terminal's data transmission module, which receives the data and passes it to the large-scale LLM.

[1030] Generating high-quality output with large-scale LLM

[1031] The server-side large-scale LLM analyzes the preprocessed queries and generates high-quality output that meets the user's requirements. For example, it generates a list of "highly rated restaurants within a 2km radius." The large-scale LLM collects and organizes information using databases and external APIs as needed.

[1032] Send and view high-quality output

[1033] The generated high-quality output is transmitted to the terminal by the server's data transmission module, and the terminal displays the received data to the user using the display module. This allows the user to quickly obtain high-quality information according to their emotions while protecting their privacy.

[1034] Specific examples

[1035] As a specific example, the following scenario is assumed.

[1036] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[1037] The device receives the input and passes it to the emotion engine.

[1038] 2. The emotion engine recognizes the user's emotions.

[1039] For example, it is determined that the user is in a hurry.

[1040] 3. A small-scale LLM performs preprocessing based on emotion information.

[1041] It clarifies the vague term "nearby" by masking location information to city level, and sets priorities for faster processing based on feelings of urgency.

[1042] 4. The preprocessed data is sent to the server.

[1043] The data transmission module of the terminal transmits the preprocessed data to the server.

[1044] 5. The server's large-scale LLM generates high-quality restaurant lists.

[1045] The server uses a database or external API to collect information about nearby highly rated restaurants.

[1046] 6. The generated list is sent to the terminal and displayed to the user.

[1047] The terminal receives the list and visually presents it to the user.

[1048] In this way, the system understands users' emotions while providing high-quality information with improved energy efficiency and privacy protection.

[1049] The processing flow will be explained below.

[1050] Step 1:

[1051] A user types a question into a smartphone app: "What are some recommended restaurants nearby?"

[1052] Step 2:

[1053] The terminal receives the user's input at a data receiving module.

[1054] Step 3:

[1055] The device sends the received input to an emotion engine that analyzes the user's emotions, such as identifying emotions like "hurried" or "relaxed" based on the user's writing style and tone.

[1056] Step 4:

[1057] If the device receives the analysis results from the emotion engine and determines that the user is in a hurry, it sends the input data to a small-scale LLM and instructs it to perform high-priority processing.

[1058] Step 5:

[1059] A small LLM analyzes user input and converts vague expressions into specific ones, for example, changing "nearby" to "within a 2km radius" and masking the user's current location to city level.

[1060] Step 6:

[1061] The terminal transmits the pre-processed data to the server via a data transmission module.

[1062] Step 7:

[1063] The server receives the preprocessed data transmitted from the terminal at a data receiving module.

[1064] Step 8:

[1065] The server passes the received data to a large-scale LLM, which analyzes the data and generates high-quality output based on the user's sentiment and requirements. For example, it creates a list of highly rated restaurants within a 2km radius.

[1066] Step 9:

[1067] The high-quality output generated by the server is transmitted to the terminal by a data transmission module.

[1068] Step 10:

[1069] The terminal receives the high-quality output transmitted from the server at a data receiving module.

[1070] Step 11:

[1071] The data received by the terminal is visually displayed to the user in a display module, for example, a list of highly rated restaurants is displayed on the screen.

[1072] Through this series of steps, users can quickly obtain high-quality information that corresponds to their emotions while their privacy is protected.

[1073] Example 2

[1074] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1075] Current language model systems perform uniform processing without considering the user's emotions, resulting in a poor user experience. They also sometimes fail to adequately protect the user's privacy and can be slow to respond. Furthermore, they lack the ability to properly concretize ambiguous input, making it difficult to provide the accurate information the user desires.

[1076] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for capturing user input, means for analyzing and preprocessing the user input using a small-scale language model, means for performing sentiment analysis, means for transmitting the preprocessed data to a large-scale language model on the cloud, means for receiving high-quality output generated by the large-scale language model, means for displaying the received high-quality output to the user, and means for concretizing ambiguous expressions based on the user input and setting priorities if necessary. This enables the provision of high-quality information that takes user emotions into consideration, protection of privacy, and highly efficient data processing.

[1077] A "small language model" is a computer program that analyzes and preprocesses user input data, optimized to operate in environments with limited computing resources.

[1078] A "large-scale language model" is a computer program that uses large amounts of data and computing resources to perform advanced natural language processing and generate high-quality output.

[1079] "Preprocessing" refers to a series of operations that analyze data entered by a user and perform masking to concretize ambiguous expressions and protect privacy.

[1080] "Receiving" means receiving data sent from a sender.

[1081] "Emotion analysis" is the process of inferring a user's emotions and intentions based on the data they input.

[1082] "Data transmission" means sending pre-processed data to a designated destination.

[1083] "High-quality output" is information generated by a large-scale language model to accurately respond to user requests.

[1084] A "data receiving module" is hardware or software that has the function of receiving data from the outside and passing it on to internal processing.

[1085] A "data transmission module" is hardware or software that has the function of transmitting pre-processed or generated data to the outside.

[1086] A "display module" is hardware or software that has the function of visually displaying received information to a user.

[1087] This system combines small-scale and large-scale language models and adds an emotion analysis engine to recognize user emotions. It efficiently processes user input and provides high-quality output that reflects the user's emotions. This system aims to protect privacy, improve energy efficiency, and respond to user emotions.

[1088] System configuration

[1089] The system is accessed by users through specific input devices (such as smartphones or tablets) and includes the following components:

[1090] Terminal

[1091] Small language models

[1092] It is a computer program that analyzes and preprocesses user input data.

[1093] Sentiment Analysis Engine

[1094] It is a computer program that analyzes emotions and intentions based on user input data.

[1095] Data Receiving Module

[1096] It has the function of receiving input from the user and passing it on to internal processing.

[1097] Data Transmission Module

[1098] It has a function for sending preprocessed data to the server.

[1099] Display Module

[1100] It has the ability to visually display high quality output received from the server to the user.

[1101] server

[1102] Large-scale language models

[1103] It is a computer program that performs advanced natural language processing and generates high-quality output that meets user requirements.

[1104] Data Receiving Module

[1105] It has the function of receiving data sent from the terminal and passing it on to internal processing.

[1106] Information Collection Module

[1107] It has the ability to collect information from databases and external APIs as needed and organize the data.

[1108] Specific examples

[1109] When a user asks a question

[1110] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[1111] The terminal's data receiving module captures this input.

[1112] 2. The sentiment analysis engine recognizes the user's emotions.

[1113] For example, the style and tone of the question can determine whether the user is in a hurry or relaxed.

[1114] 3. A small language model performs preprocessing based on emotion information.

[1115] The vague expression "nearby" is specified as "within a 2km radius," location information is masked to the city level, and priority is set for fast processing based on the feeling of urgency.

[1116] 4. The preprocessed data is sent to the server.

[1117] The data transmission module of the terminal transmits the preprocessed data to the server.

[1118] 5. The server's large-scale language model generates high-quality restaurant lists.

[1119] A large language model uses databases and external APIs to create a list of "highly rated restaurants within a 2km radius."

[1120] 6. The generated list is sent to the terminal and displayed to the user.

[1121] The terminal receives the list and visually presents it to the user using a display module.

[1122] Examples of prompt statements

[1123] For example, if a user inputs a question such as "What are some recommended restaurants nearby?", the above process will promptly provide high-quality information while taking into consideration the user's feelings.

[1124] In this way, this system understands the user's emotions and provides high-quality information while protecting privacy.

[1125] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1126] Step 1: Capturing User Input

[1127] A user types into their smartphone, "What are some recommended restaurants near me?"

[1128] The terminal's data reception module captures this input. Specifically, it acquires the string entered in the text box as internal data.

[1129] Input: User question (e.g., "What are some recommended restaurants nearby?")

[1130] Output: Captured user question data (e.g., "What are some recommended restaurants nearby?")

[1131] Step 2: Emotion recognition by the emotion engine

[1132] The device sends the captured user question to the emotion engine, which analyzes emotions from the input text.

[1133] Input: User question data (e.g., "What are some recommended restaurants nearby?")

[1134] Output: Emotional information (e.g., user is in a hurry)

[1135] Step 3: Preprocessing with small-scale LLM

[1136] The device passes the captured question data based on the emotion information to the small-scale LLM, which analyzes the question and makes ambiguous expressions concrete.

[1137] Input: User question data and emotion information (e.g., "What are some recommended restaurants nearby?", emotion: "I'm in a hurry")

[1138] Specific behavior: Converts the vague expression "nearby" to "within a 2km radius," masks location information to the city / ward / town / village level for privacy reasons, and sets priorities for faster processing.

[1139] Output: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", sentiment: in a hurry)

[1140] Step 4: Sending preprocessed data

[1141] The data transmission module of the terminal transmits the preprocessed data to the server, specifically, by using an HTTP request.

[1142] Input: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", Sentiment: Hurry)

[1143] Output: Data sent to the server

[1144] Step 5: Generating high-quality output with large-scale LLM

[1145] The server receives the pre-processed data and passes it to the large-scale LLM, which uses databases and external APIs to generate high-quality output.

[1146] Input: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", Sentiment: Hurry)

[1147] Specific behavior: Performs database queries and external API calls, and generates a list of restaurants based on the obtained data.

[1148] Output: High-quality output data (e.g., "List of highly rated restaurants within a 2km radius")

[1149] Step 6: Send and view high-quality output

[1150] The data transmission module of the server transmits the generated high-quality output to the terminal.

[1151] The terminal receives the output sent from the server and visually displays it to the user using a display module.

[1152] Input: High-quality output data (e.g., "List of highly rated restaurants within a 2km radius")

[1153] Specific behavior: Renders the received data in the UI and displays it to the user.

[1154] Output: High-quality information visually presented to the user

[1155] In this way, the system understands users' emotions while providing high-quality information with improved energy efficiency and privacy protection.

[1156] (Application example 2)

[1157] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1158] Conventional systems have had difficulty in accurately recognizing emotions in response to user input and providing personalized, high-quality information based on that emotion. Furthermore, while there is a need to efficiently provide information tailored to the user's emotions, there are insufficient measures to protect privacy and prevent unnecessary repetition of questions. Understanding customer emotions and providing real-time services based on those emotions is particularly important in brick-and-mortar stores, but conventional technologies have had difficulty meeting this requirement.

[1159] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing and preprocessing a user's input using a small-scale language model, means for transmitting the preprocessed data to a large-scale language model on the cloud, means for receiving high-quality output generated by the large-scale language model, means for displaying the received high-quality output to the user, means for recognizing emotions from the user's input, and means for adjusting data preprocessing based on the recognized emotions. This makes it possible to provide information according to the user's emotions in real time, protect privacy, and prevent unnecessary repetition of questions.

[1160] A "small language model" is a lightweight, energy-efficient natural language processing model used to parse and preprocess user input.

[1161] A "large-scale language model" is a high-performance natural language processing model that runs on the cloud and generates high-quality output based on pre-processed data.

[1162] "User input" refers to data such as text information and voice information that a user provides to the system.

[1163] "Preprocessing" refers to the initial processing of user input to parse it and convert it into a format that meets specific needs.

[1164] "Sending to the cloud" refers to transferring data from the user's terminal to a remote server.

[1165] "High-quality output" refers to highly accurate information and results generated by a large-scale language model that are tailored to the user's needs.

[1166] "Receiving" refers to receiving the generated output from the server at the user's terminal.

[1167] "Displaying" refers to visually presenting the received high quality output to a user.

[1168] "Emotion recognition" refers to the process of determining a user's emotional state from their input.

[1169] "Adjusting data preprocessing" refers to changing and optimizing the content and methods of preprocessing based on the recognized emotions.

[1170] The system for implementing this invention analyzes user input, recognizes emotions, and provides high-quality output based on those emotions. The system consists of two main components: a terminal and a server.

[1171] Program Overview

[1172] 1. Capturing user input

[1173] The user inputs data through a terminal (e.g., a smartphone). The input data can be text data or voice data.

[1174] 2. Emotional Recognition

[1175] The device passes the input data to an emotion engine (e.g., Hugging Face sentiment-analysis) to recognize the user's emotion. For example, if the user input is "Are there any deals?", the device recognizes the user's emotion as "happy" based on that input.

[1176] 3. Pretreatment

[1177] It uses a small language model (e.g., T5-small) to analyze user input and convert it into a concrete form while taking into account emotional information, for example, converting the vague expression "great deal" into "today's sale items," and adjusts the processing speed based on the emotion.

[1178] 4. Data transmission

[1179] The pre-processed data is transmitted to the server by the data transmission module of the terminal.

[1180] 5. Processing with large-scale language models

[1181] The server uses large-scale language models to analyze the pre-processed data and generate high-quality output, such as a list of today's sale items.

[1182] 6. Send and display high-quality output

[1183] The server generates high-quality output and sends it to the terminal, which visually displays the received data to the user.

[1184] Hardware and software used

[1185] Device: User device such as a smartphone or tablet

[1186] Emotion Engine: A sentiment-analysis model of Hugging Face

[1187] Small language models: Natural language processing models such as T5-small

[1188] Server: High-performance server on the cloud

[1189] Large-scale language model: High-performance natural language processing model running on a server

[1190] Specific examples

[1191] The specific scenario is as follows:

[1192] 1. The user speaks into the smartphone app and says, "Do you have any deals?"

[1193] 2. The device captures the input and recognizes the "happy" emotion using its emotion engine.

[1194] 3. Small LLMs concretize the expression "good deal" and transform it into "What's on sale today?"

[1195] 4. Send it to the server, and the large-scale LLM generates a "list of today's sale items" that meets your request.

[1196] 5. Once the list is returned, the device displays it to the user along with "Today's Sale Items."

[1197] Example prompt sentence:

[1198] Q: Are there any great deals? I'd be happy to provide information at a speed that matches my needs. A:

[1199] In this way, the system can provide personalized information based on user sentiment in real time, improving the customer experience.

[1200] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1201] Step 1:

[1202] The user types "Do you have any deals?" into the smartphone app. The input can be text or voice. The device captures this data and stores it as input data.

[1203] Step 2:

[1204] The device passes the captured input data to the emotion engine, which (Hugging Face's sentiment-analysis model) recognizes the emotion from the user's input. In this step, based on the input data (e.g., "Are there any deals available?"), the emotion engine outputs the emotion information "happy."

[1205] Step 3:

[1206] The device takes the recognized emotional information into consideration and passes the user's input to a small-scale LLM (T5-small model). The small-scale LLM analyzes the input and performs data preprocessing based on the emotional information. In this step, the device outputs preprocessed data (e.g., "What are the sale items today?") based on the input data and emotional information (e.g., "Are there any deals?" and "I'm happy").

[1207] Step 4:

[1208] The data transmission module of the device sends the preprocessed data to a server on the cloud, where the preprocessed data ("What are the sale items today?") is transferred to the server as input data.

[1209] Step 5:

[1210] The server passes the received preprocessed data to the large-scale LLM, which analyzes the preprocessed data and generates the corresponding high-quality output. In this step, the server outputs a high-quality output ("List of sale items") based on the preprocessed input data ("What items are on sale today?").

[1211] Step 6:

[1212] The server's data transmission module transmits the generated high-quality output to the terminal, where the generated data ("list of sale items") is transmitted to the terminal.

[1213] Step 7:

[1214] The terminal uses a display module to display the received high-quality data to the user. In this step, the terminal visually presents the high-quality output data ("list of sale items") to the user.

[1215] The above is the specific process flow from user input to the provision of high-quality information. Data is input and output at each step, and the next step is executed based on the results, thereby realizing the provision of highly accurate, emotion-responsive information.

[1216] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1217] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1218] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1219] [Fourth embodiment]

[1220] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1221] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1222] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1223] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1224] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1226] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1227] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1228] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1229] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1230] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1231] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1232] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1233] The present invention is a system that combines small-scale language models (hereinafter referred to as small-scale LLMs) and large-scale language models (hereinafter referred to as large-scale LLMs), which enables efficient processing of user input and provides high-quality output. This system aims to protect privacy and improve energy efficiency, among other things.

[1234] System configuration

[1235] This system is intended to be used by users through devices such as smartphones and tablets. The system includes the following components:

[1236] Terminal

[1237] Small scale LLM

[1238] Data reception and transmission module

[1239] Display Module

[1240] server

[1241] large scale LLM

[1242] Data reception and transmission module

[1243] Information Collection Module

[1244] System Operation

[1245] When a user uses the system to ask a question, the following process is carried out.

[1246] Capturing User Input

[1247] A user inputs a question into a smartphone app. For example, they input a question like, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[1248] Pretreatment by small-scale LLM

[1249] The device first sends the captured input to a small-scale LLM for preprocessing. Specifically, it removes ambiguous expressions from the question and masks the user's privacy information (e.g., current location information). For example, it clarifies the expression "nearby" to "within a 2km radius." This preprocessing aims to clarify the question while protecting the user's privacy.

[1250] Sending preprocessed data

[1251] The pre-processed data is sent to the server by the terminal's data transmission module, which receives the data and passes it to the large-scale LLM.

[1252] Generating high-quality output with large-scale LLM

[1253] The server-side large-scale LLM analyzes the preprocessed queries and generates high-quality output that meets the user's requirements. For example, it generates a list of "highly rated restaurants within a 2km radius." The large-scale LLM collects and organizes information using databases and external APIs as needed.

[1254] Send and view high-quality output

[1255] The generated high-quality output is sent to the terminal by the server's data transmission module, and the terminal displays the received data to the user using the display module, allowing the user to accurately and quickly obtain the information they desire while protecting their privacy.

[1256] Specific examples

[1257] As a specific example, the following scenario is assumed.

[1258] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[1259] The terminal receives the input and passes it to the small-scale LLM.

[1260] 2. A small LLM preprocesses the user input.

[1261] For example, the vague expression "nearby" is made more specific, and the location information is masked to the city / ward / town / village level.

[1262] 3. The preprocessed data is sent to the server.

[1263] The data transmission module of the terminal transmits the preprocessed data to the server.

[1264] 4. The server's large-scale LLM generates high-quality restaurant lists.

[1265] The server uses a database or external API to retrieve information about nearby highly rated restaurants.

[1266] 5. The generated list is sent to the terminal and displayed to the user.

[1267] The terminal receives the list and visually presents it to the user.

[1268] In this way, the system provides high-quality, energy-efficient information while protecting user privacy.

[1269] The processing flow will be explained below.

[1270] Step 1:

[1271] A user types a question into a smartphone app: "What are some recommended restaurants nearby?"

[1272] Step 2:

[1273] The terminal receives the user's input at a data receiving module.

[1274] Step 3:

[1275] The device passes the received input to the small-scale LLM, which analyzes the question and translates it into a more specific expression. For example, it converts "nearby" to "within a 2km radius" and masks the user's current location to the city level.

[1276] Step 4:

[1277] The terminal transmits the pre-processed data to the server via a data transmission module.

[1278] Step 5:

[1279] The server receives the preprocessed data transmitted from the terminal at a data receiving module.

[1280] Step 6:

[1281] The server passes the received data to a large-scale LLM, which analyzes the user's question and generates high-quality output. For example, it may use a database or external API to collect information on highly rated restaurants within a 2km radius.

[1282] Step 7:

[1283] The server transmits high-quality output to the terminal via a data transmission module.

[1284] Step 8:

[1285] The terminal receives the high-quality output transmitted from the server at a data receiving module.

[1286] Step 9:

[1287] The data received by the terminal is visually displayed to the user in a display module, for example, a list of highly rated restaurants is displayed on the screen.

[1288] This series of processes allows users to quickly obtain high-quality information that protects their privacy in an energy-efficient manner.

[1289] Example 1

[1290] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1291] Modern information processing systems are required to quickly and accurately provide high-quality output in response to user-entered questions. However, achieving this goal poses several challenges. First, from the perspective of privacy protection, users' personal information must be handled securely. Second, if the input question is ambiguous, preprocessing is required to make it more specific. Furthermore, generating high-quality output requires processing large amounts of data, and improving energy efficiency for this purpose is also an important challenge.

[1292] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1293] In this invention, the server includes means for analyzing and preprocessing user input using a small-scale language model installed on the user terminal, means for transmitting the preprocessed data to a large-scale language model on the cloud, and means for receiving high-quality output generated by the large-scale language model, thereby enabling ambiguous questions to be specified and providing high-quality information with high energy efficiency while protecting user privacy.

[1294] A "small language model" is a language processing algorithm that runs on a terminal and analyzes and preprocesses user input.

[1295] "Preprocessing" refers to the process of removing ambiguous expressions from input data and anonymizing user privacy information as necessary.

[1296] "Anonymization" refers to making a user's information in a state where it is not possible to identify the individual in order to protect the user's privacy information.

[1297] "Concreteization" is the process of converting vague expressions into clear, specific expressions.

[1298] A "large-scale language model" is an advanced language processing algorithm that runs on the cloud and produces high-quality output.

[1299] "High-quality output" refers to relevant and accurate answers and information to users' questions.

[1300] A "database" is a system for systematically storing and managing data.

[1301] An "external API" is a program interface that connects to external databases and services to obtain information.

[1302] A "secure communication protocol" is a communication rule for safely sending and receiving data.

[1303] A "user terminal" is a computer device that can be directly operated by a user, such as a smartphone or tablet.

[1304] "Analysis" is the act of breaking down and examining input data in order to understand it and perform the necessary processing.

[1305] "Masking" is the act of hiding or making certain information invisible.

[1306] The present invention is an information processing system that combines a small-scale language model installed on a user terminal with a large-scale language model on the cloud. The system aims to provide high-quality output in response to user input. Specifically, the small-scale language model preprocesses data entered by the user, and the preprocessed data is sent to a large-scale language model on the cloud for analysis. The system then receives the high-quality output generated by the large-scale language model and displays it to the user.

[1307] Hardware and software used

[1308] User device: A computing device that can be directly operated by a user, such as a smartphone or tablet. The device is equipped with a small language model, a data reception / transmission module, and a display module.

[1309] Cloud server: A server on which a large-scale language model runs, and includes a data reception / transmission module and an information collection module.

[1310] Small language model behavior

[1311] A small language model installed on the user's device analyzes data entered by the user (e.g., "What are some recommended restaurants nearby?") and performs preprocessing. This preprocessing involves concretizing ambiguous expressions and anonymizing private information. For example, the ambiguous expression "nearby" is concretized as "within a 2km radius," and location information is converted to the city, town, or village level.

[1312] Sending data

[1313] The pre-processed data is then sent to the cloud server via the device's data transmission module, where data security is ensured using secure communication protocols such as HTTPS.

[1314] How large language models work

[1315] The preprocessed data sent to the cloud server is analyzed by a large-scale language model, which gathers information from databases and external APIs to generate the best answer to the user's question, such as a list of "highly rated restaurants within a 2km radius."

[1316] Send and view high-quality output

[1317] The generated high-quality output is sent to the user terminal by the server's data transmission module, and the terminal visually presents the received data to the user via the display module, allowing the user to obtain the information they need accurately and quickly while protecting their privacy.

[1318] Specific examples

[1319] 1. User Input: A user types into their smartphone, "What are some recommended restaurants nearby?"

[1320] 2. Preprocessing: A small language model concretizes the expression "nearby" to "within a 2km radius" and anonymizes location information to the city / ward / town / village level.

[1321] 3. Data transmission: The preprocessed data is transmitted to the cloud server using HTTPS.

[1322] 4. Parsing and output generation: A large language model uses databases and external APIs to create a list of "top rated restaurants within a 2km radius."

[1323] 5. Sending and displaying the output: The generated list is sent to the user terminal and presented to the user via the display module.

[1324] In this way, the present invention provides high-quality information while protecting the user's privacy.

[1325] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1326] Step 1: Capturing User Input

[1327] A user types "What are some recommended restaurants nearby?" into a smartphone app. The device captures this input in a data receiving module and passes it to a small language model as text data. At this stage, the input data is the raw text entered by the user.

[1328] Step 2: Preprocessing with small-scale LLM

[1329] A small language model on the device receives the captured input text and analyzes it. For the input text "What are some recommended restaurants nearby?", the ambiguous "nearby" is made more specific to "within a 2km radius," and the user's location information is further anonymized at the city / ward / town / village level. The output of this step is the preprocessed text data "What are some recommended restaurants within a 2km radius?"

[1330] Step 3: Sending preprocessed data

[1331] The preprocessed data is received by the terminal's data transmission module and sent to the cloud server. HTTPS is used for this transmission to ensure data security. The input is the preprocessed text data, and the output is the encoded data transmission.

[1332] Step 4: Receiving data on the server

[1333] The server receives data sent from the device through a data reception module, decodes the received encoded data, and prepares it for passing to a large-scale language model, where the input is the encoded preprocessed data and the output is the decoded preprocessed data.

[1334] Step 5: Generating high-quality output with large-scale LLM

[1335] The server's large-scale language model analyzes the preprocessed data. Specifically, for the question "What are some recommended restaurants within a 2km radius?", it collects information from databases and external APIs and generates a list of "highly rated restaurants within a 2km radius." The input in this step is the decoded preprocessed data, and the output is a high-quality restaurant list.

[1336] Step 6: Send high-quality output

[1337] The generated high-quality output is encoded by the server's data transmission module and sent to the terminal. Again, HTTPS is used to transmit data securely. The input is a high-quality restaurant list, and the output is the encoded restaurant list.

[1338] Step 7: Receive and display data on your device

[1339] The terminal's data receiving module receives high-quality output from the server. The received data is decoded and passed to the display module. The display module displays the information in a visually easy-to-understand format for the user, for example, as a "list of highly rated restaurants." The input in this step is the encoded restaurant list, and the output is restaurant information visually displayed to the user.

[1340] This allows users to obtain the information they need accurately and quickly while their privacy is protected.

[1341] (Application example 1)

[1342] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1343] Conventional language modeling systems have problems such as insufficient protection of user privacy information and difficulty in obtaining high-quality answers due to ambiguous user questions. Energy efficiency is also an issue. Food delivery services, in particular, require users to quickly and accurately find nearby restaurant recommendations.

[1344] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1345] In this invention, the server includes: means for analyzing and preprocessing user input using a small-scale language model; means for sending the preprocessed data to a large-scale language model on the cloud; means for receiving high-quality output generated by the large-scale language model; means for displaying the received high-quality output to the user; means for converting ambiguous expressions into specific conditions and anonymizing location information in the preprocessing; means for masking privacy information to the city / ward / town / village level to protect user privacy; means for generating a restaurant list in descending order of rating; means for collecting and organizing information based on the preprocessed data; and means for generating high-quality output from prompt sentences using a generative AI model. This enables ambiguous questions to be converted into specific, high-quality answers while protecting user privacy. It also enables users to quickly and accurately find recommended nearby restaurants for food delivery.

[1346] A "small language model" is a relatively small language processing system capable of analyzing user input and converting vague expressions into concrete terms.

[1347] "Preprocessing" is the process of analyzing user input information, concretizing ambiguous expressions, and anonymizing and masking privacy information.

[1348] A "large-scale language model in the cloud" is a large-scale language processing system accessed via the internet for complex analysis and high-quality output generation.

[1349] "High-quality output" is accurate and detailed information generated by a large-scale language model and tailored to the user's needs.

[1350] "Means for displaying to a user" refers to a display or display module for visually presenting the received output on a terminal.

[1351] "Ambiguous expressions" refer to parts of the user's input that lack specificity or are unclear.

[1352] "Specific conditions" are those that convert vague expressions into clear and detailed information.

[1353] "Means for anonymizing location information" refers to a method for protecting privacy by deleting or converting the user's current location to a level that makes it identifiable.

[1354] "Means for masking privacy information to the city, ward, town, or village level" refers to a method for preventing individuals from being identified by limiting the user's location information to the city, ward, town, or village level.

[1355] The "means for generating a restaurant list in descending order of rating" is a method for sorting nearby restaurants in a food delivery service by rating criteria.

[1356] "Means for collecting and organizing information based on preprocessed data" refers to a method for obtaining necessary information from preprocessed data and organizing it appropriately.

[1357] A "generative AI model" is an artificial intelligence language processing model that provides appropriate output based on an input prompt.

[1358] A "prompt" is an instruction or question given to an AI model to generate a specific output.

[1359] The system for implementing this invention consists of a terminal that analyzes and preprocesses user input using a small-scale language model, and a server that generates high-quality output using a large-scale language model on the cloud.

[1360] System configuration

[1361] The system consists of the following hardware and software:

[1362] Hardware

[1363] User devices (smartphones, tablets)

[1364] Cloud server with graphics card

[1365] software

[1366] Small language model (runs on device)

[1367] Large-scale language model (running on cloud servers)

[1368] Transformers Library (Hugging Face)

[1369] requests library (data communication)

[1370] React Native (user interface)

[1371] System Operation

[1372] The system operates in the following manner to efficiently process user input information and provide high quality output.

[1373] 1. Capturing User Input

[1374] A user uses a smartphone app to input a question, for example, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[1375] 2. Pretreatment by small-scale LLM

[1376] User input is first sent to a small language model on the device for preprocessing. For example, a vague expression like "nearby" is converted into a more specific condition like "within a 2km radius." Additionally, user location information is anonymized at the city / ward / town / village level, protecting user privacy.

[1377] 3. Sending preprocessed data

[1378] After the preprocessing is complete, the data is sent to the cloud server by the terminal's data transmission module, which receives the data and passes it to a large-scale language model.

[1379] 4. Generating high-quality output using large-scale LLM

[1380] A large-scale language model on a cloud server analyzes the preprocessed questions and generates high-quality output, such as a list of "top-rated restaurants within a 2km radius." The large-scale language model collects and organizes information using databases and external APIs as needed.

[1381] 5. Send and display high-quality output

[1382] The generated high-quality output is sent to the device by the cloud server's data transmission module, and the device displays the received data to the user via the display module, allowing the user to accurately and quickly obtain the information they desire while protecting their privacy.

[1383] Examples of concrete examples and prompts

[1384] A concrete example of how the system has been implemented is a food delivery application, which works as follows:

[1385] 1. A user types into the app, "What are some recommended restaurants nearby?"

[1386] 2. The app converts the input into "What are some recommended restaurants within a 2km radius?" and anonymizes the location information to "Shinjuku-ku, Tokyo" before sending it to the server.

[1387] 3. The server uses an AI model to generate a list of highly rated restaurants nearby.

[1388] 4. Send the list to the app and display it to the user.

[1389] Example prompt sentence:

[1390] Please list recommended restaurants within a 2km radius. The user's anonymized location is "Shinjuku-ku, Tokyo."

[1391] In this way, it is possible to provide highly accurate restaurant information for food delivery while ensuring the user's privacy.

[1392] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1393] Step 1:

[1394] A user opens a smartphone application and inputs the question, "What are some recommended restaurants nearby?" This input is captured on the user's device and passed to a data receiving module. The input data is natural language text, and the operation to capture the input data is performed via a user interface. In this case, the input is "What are some recommended restaurants nearby?" and the output is the user's input text data.

[1395] Step 2:

[1396] The device passes the captured user input to a small-scale language model and begins preprocessing. Specifically, the small-scale LLM concretizes the ambiguous expression "nearby" to "within a 2km radius" and anonymizes the user's location information to "Shinjuku Ward, Tokyo" at the city / ward / town / village level. The input for preprocessing is the captured user's natural language text and location information, and the output is a preprocessed concretized question: "Recommend restaurants within a 2km radius (Shinjuku Ward, Tokyo)."

[1397] Step 3:

[1398] The preprocessed data is sent to the cloud server by the terminal's data transmission module. During the transmission process, the data is encrypted using a protocol such as HTTPS request and sent to the server. The input is the preprocessed query data, and the output is the completion of data transmission to the cloud server.

[1399] Step 4:

[1400] The cloud server passes the received preprocessed data to a large-scale language model (generative AI model). The large-scale LLM uses external databases and APIs to provide high-quality information. Specifically, it generates a list of "highly rated restaurants within a 2km radius." The input of the large-scale LLM is the preprocessed, specified question data, and the output is a list of highly rated restaurants.

[1401] Step 5:

[1402] The server receives the high-quality output (restaurant list) generated by the large-scale LLM and sends it to the terminal. The transmission process is also encrypted and is carried out using protocols such as HTTPS requests. The input here is the generated restaurant list data, and the output is the completion of transmission of the restaurant list to the terminal.

[1403] Step 6:

[1404] The terminal passes the restaurant list received from the server to the display module, which then visually presents it to the user. Specifically, the restaurant list is displayed on the application interface using a framework such as React Native. The input here is the received restaurant list data, and the output is the specific restaurant list visually displayed to the user.

[1405] This series of processing steps allows users to quickly obtain high-quality restaurant information for food delivery services while protecting their privacy.

[1406] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1407] This invention is a system that combines a small-scale language model (hereinafter referred to as small-scale LLM) with a large-scale language model (hereinafter referred to as large-scale LLM), and adds an emotion engine that recognizes user emotions. It efficiently processes user input and provides high-quality output according to the user's emotions. The system aims to protect privacy, improve energy efficiency, and respond to user emotions.

[1408] System configuration

[1409] The system is accessed by users through specific input devices (such as smartphones or tablets) and includes the following components:

[1410] Terminal

[1411] Small scale LLM

[1412] Emotion Engine

[1413] Data reception and transmission module

[1414] Display Module

[1415] server

[1416] large scale LLM

[1417] Data reception and transmission module

[1418] Information Collection Module

[1419] System Operation

[1420] When a user uses the system to ask a question, the following process is carried out.

[1421] Capturing User Input

[1422] A user inputs a question into a smartphone app. For example, they input a question like, "What are some recommended restaurants nearby?" This input is captured by the device's data receiving module.

[1423] Emotion recognition with emotion engine

[1424] The device sends the captured input to an emotion engine to recognize the user's emotions, for example, determining whether the user is in a hurry or relaxed state based on the style and tone of the question.

[1425] Pretreatment by small-scale LLM

[1426] The device uses the emotion information obtained from the emotion engine and passes it to the small-scale LLM. The small-scale LLM analyzes the question and converts it into a more specific expression. For example, it converts the vague expression "nearby" into "within a 2km radius" and masks the user's current location information to the city / ward / town / village level. If the device detects a sense of urgency, it sets a priority to speed up processing.

[1427] Sending preprocessed data

[1428] The pre-processed data is sent to the server by the terminal's data transmission module, which receives the data and passes it to the large-scale LLM.

[1429] Generating high-quality output with large-scale LLM

[1430] The server-side large-scale LLM analyzes the preprocessed queries and generates high-quality output that meets the user's requirements. For example, it generates a list of "highly rated restaurants within a 2km radius." The large-scale LLM collects and organizes information using databases and external APIs as needed.

[1431] Send and view high-quality output

[1432] The generated high-quality output is transmitted to the terminal by the server's data transmission module, and the terminal displays the received data to the user using the display module. This allows the user to quickly obtain high-quality information according to their emotions while protecting their privacy.

[1433] Specific examples

[1434] As a specific example, the following scenario is assumed.

[1435] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[1436] The device receives the input and passes it to the emotion engine.

[1437] 2. The emotion engine recognizes the user's emotions.

[1438] For example, it is determined that the user is in a hurry.

[1439] 3. A small-scale LLM performs preprocessing based on emotion information.

[1440] It clarifies the vague term "nearby" by masking location information to city level, and sets priorities for faster processing based on feelings of urgency.

[1441] 4. The preprocessed data is sent to the server.

[1442] The data transmission module of the terminal transmits the preprocessed data to the server.

[1443] 5. The server's large-scale LLM generates high-quality restaurant lists.

[1444] The server uses a database or external API to collect information about nearby highly rated restaurants.

[1445] 6. The generated list is sent to the terminal and displayed to the user.

[1446] The terminal receives the list and visually presents it to the user.

[1447] In this way, the system understands users' emotions while providing high-quality information with improved energy efficiency and privacy protection.

[1448] The processing flow will be explained below.

[1449] Step 1:

[1450] A user types a question into a smartphone app: "What are some recommended restaurants nearby?"

[1451] Step 2:

[1452] The terminal receives the user's input at a data receiving module.

[1453] Step 3:

[1454] The device sends the received input to an emotion engine that analyzes the user's emotions, such as identifying emotions like "hurried" or "relaxed" based on the user's writing style and tone.

[1455] Step 4:

[1456] If the device receives the analysis results from the emotion engine and determines that the user is in a hurry, it sends the input data to a small-scale LLM and instructs it to perform high-priority processing.

[1457] Step 5:

[1458] A small LLM analyzes user input and converts vague expressions into specific ones, for example, changing "nearby" to "within a 2km radius" and masking the user's current location to city level.

[1459] Step 6:

[1460] The terminal transmits the pre-processed data to the server via a data transmission module.

[1461] Step 7:

[1462] The server receives the preprocessed data transmitted from the terminal at a data receiving module.

[1463] Step 8:

[1464] The server passes the received data to a large-scale LLM, which analyzes the data and generates high-quality output based on the user's sentiment and requirements. For example, it creates a list of highly rated restaurants within a 2km radius.

[1465] Step 9:

[1466] The high-quality output generated by the server is transmitted to the terminal by a data transmission module.

[1467] Step 10:

[1468] The terminal receives the high-quality output transmitted from the server at a data receiving module.

[1469] Step 11:

[1470] The data received by the terminal is visually displayed to the user in a display module, for example, a list of highly rated restaurants is displayed on the screen.

[1471] Through this series of steps, users can quickly obtain high-quality information that corresponds to their emotions while their privacy is protected.

[1472] Example 2

[1473] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1474] Current language model systems perform uniform processing without considering the user's emotions, resulting in a poor user experience. They also sometimes fail to adequately protect the user's privacy and can be slow to respond. Furthermore, they lack the ability to properly concretize ambiguous input, making it difficult to provide the accurate information the user desires.

[1475] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for capturing user input, means for analyzing and preprocessing the user input using a small-scale language model, means for performing sentiment analysis, means for transmitting the preprocessed data to a large-scale language model on the cloud, means for receiving high-quality output generated by the large-scale language model, means for displaying the received high-quality output to the user, and means for concretizing ambiguous expressions based on the user input and setting priorities if necessary. This enables the provision of high-quality information that takes user emotions into consideration, protection of privacy, and highly efficient data processing.

[1476] A "small language model" is a computer program that analyzes and preprocesses user input data, optimized to operate in environments with limited computing resources.

[1477] A "large-scale language model" is a computer program that uses large amounts of data and computing resources to perform advanced natural language processing and generate high-quality output.

[1478] "Preprocessing" refers to a series of operations that analyze data entered by a user and perform masking to concretize ambiguous expressions and protect privacy.

[1479] "Receiving" means receiving data sent from a sender.

[1480] "Emotion analysis" is the process of inferring a user's emotions and intentions based on the data they input.

[1481] "Data transmission" means sending pre-processed data to a designated destination.

[1482] "High-quality output" is information generated by a large-scale language model to accurately respond to user requests.

[1483] A "data receiving module" is hardware or software that has the function of receiving data from the outside and passing it on to internal processing.

[1484] A "data transmission module" is hardware or software that has the function of transmitting pre-processed or generated data to the outside.

[1485] A "display module" is hardware or software that has the function of visually displaying received information to a user.

[1486] This system combines small-scale and large-scale language models and adds an emotion analysis engine to recognize user emotions. It efficiently processes user input and provides high-quality output that reflects the user's emotions. This system aims to protect privacy, improve energy efficiency, and respond to user emotions.

[1487] System configuration

[1488] The system is accessed by users through specific input devices (such as smartphones or tablets) and includes the following components:

[1489] Terminal

[1490] Small language models

[1491] It is a computer program that analyzes and preprocesses user input data.

[1492] Sentiment Analysis Engine

[1493] It is a computer program that analyzes emotions and intentions based on user input data.

[1494] Data Receiving Module

[1495] It has the function of receiving input from the user and passing it on to internal processing.

[1496] Data Transmission Module

[1497] It has a function for sending preprocessed data to the server.

[1498] Display Module

[1499] It has the ability to visually display high quality output received from the server to the user.

[1500] server

[1501] Large-scale language models

[1502] It is a computer program that performs advanced natural language processing and generates high-quality output that meets user requirements.

[1503] Data Receiving Module

[1504] It has the function of receiving data sent from the terminal and passing it on to internal processing.

[1505] Information Collection Module

[1506] It has the ability to collect information from databases and external APIs as needed and organize the data.

[1507] Specific examples

[1508] When a user asks a question

[1509] 1. A user types into their smartphone, "What are some recommended restaurants near me?"

[1510] The terminal's data receiving module captures this input.

[1511] 2. The sentiment analysis engine recognizes the user's emotions.

[1512] For example, the style and tone of the question can determine whether the user is in a hurry or relaxed.

[1513] 3. A small language model performs preprocessing based on emotion information.

[1514] The vague expression "nearby" is specified as "within a 2km radius," location information is masked to the city level, and priority is set for fast processing based on the feeling of urgency.

[1515] 4. The preprocessed data is sent to the server.

[1516] The data transmission module of the terminal transmits the preprocessed data to the server.

[1517] 5. The server's large-scale language model generates high-quality restaurant lists.

[1518] A large language model uses databases and external APIs to create a list of "highly rated restaurants within a 2km radius."

[1519] 6. The generated list is sent to the terminal and displayed to the user.

[1520] The terminal receives the list and visually presents it to the user using a display module.

[1521] Examples of prompt statements

[1522] For example, if a user inputs a question such as "What are some recommended restaurants nearby?", the above process will promptly provide high-quality information while taking into consideration the user's feelings.

[1523] In this way, this system understands the user's emotions and provides high-quality information while protecting privacy.

[1524] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1525] Step 1: Capturing User Input

[1526] A user types into their smartphone, "What are some recommended restaurants near me?"

[1527] The terminal's data reception module captures this input. Specifically, it acquires the string entered in the text box as internal data.

[1528] Input: User question (e.g., "What are some recommended restaurants nearby?")

[1529] Output: Captured user question data (e.g., "What are some recommended restaurants nearby?")

[1530] Step 2: Emotion recognition by the emotion engine

[1531] The device sends the captured user question to the emotion engine, which analyzes emotions from the input text.

[1532] Input: User question data (e.g., "What are some recommended restaurants nearby?")

[1533] Output: Emotional information (e.g., user is in a hurry)

[1534] Step 3: Preprocessing with small-scale LLM

[1535] The device passes the captured question data based on the emotion information to the small-scale LLM, which analyzes the question and makes ambiguous expressions concrete.

[1536] Input: User question data and emotion information (e.g., "What are some recommended restaurants nearby?", emotion: "I'm in a hurry")

[1537] Specific behavior: Converts the vague expression "nearby" to "within a 2km radius," masks location information to the city / ward / town / village level for privacy reasons, and sets priorities for faster processing.

[1538] Output: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", sentiment: in a hurry)

[1539] Step 4: Sending preprocessed data

[1540] The data transmission module of the terminal transmits the preprocessed data to the server, specifically, by using an HTTP request.

[1541] Input: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", Sentiment: Hurry)

[1542] Output: Data sent to the server

[1543] Step 5: Generating high-quality output with large-scale LLM

[1544] The server receives the pre-processed data and passes it to the large-scale LLM, which uses databases and external APIs to generate high-quality output.

[1545] Input: Preprocessed question data (e.g., "What are some recommended restaurants within a 2km radius?", Sentiment: Hurry)

[1546] Specific behavior: Performs database queries and external API calls, and generates a list of restaurants based on the obtained data.

[1547] Output: High-quality output data (e.g., "List of highly rated restaurants within a 2km radius")

[1548] Step 6: Send and view high-quality output

[1549] The data transmission module of the server transmits the generated high-quality output to the terminal.

[1550] The terminal receives the output sent from the server and visually displays it to the user using a display module.

[1551] Input: High-quality output data (e.g., "List of highly rated restaurants within a 2km radius")

[1552] Specific behavior: Renders the received data in the UI and displays it to the user.

[1553] Output: High-quality information visually presented to the user

[1554] In this way, the system understands users' emotions while providing high-quality information with improved energy efficiency and privacy protection.

[1555] (Application example 2)

[1556] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1557] Conventional systems have had difficulty in accurately recognizing emotions in response to user input and providing personalized, high-quality information based on that emotion. Furthermore, while there is a need to efficiently provide information tailored to the user's emotions, there are insufficient measures to protect privacy and prevent unnecessary repetition of questions. Understanding customer emotions and providing real-time services based on those emotions is particularly important in brick-and-mortar stores, but conventional technologies have had difficulty meeting this requirement.

[1558] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing and preprocessing a user's input using a small-scale language model, means for transmitting the preprocessed data to a large-scale language model on the cloud, means for receiving high-quality output generated by the large-scale language model, means for displaying the received high-quality output to the user, means for recognizing emotions from the user's input, and means for adjusting data preprocessing based on the recognized emotions. This makes it possible to provide information according to the user's emotions in real time, protect privacy, and prevent unnecessary repetition of questions.

[1559] A "small language model" is a lightweight, energy-efficient natural language processing model used to parse and preprocess user input.

[1560] A "large-scale language model" is a high-performance natural language processing model that runs on the cloud and generates high-quality output based on pre-processed data.

[1561] "User input" refers to data such as text information and voice information that a user provides to the system.

[1562] "Preprocessing" refers to the initial processing of user input to parse it and convert it into a format that meets specific needs.

[1563] "Sending to the cloud" refers to transferring data from the user's terminal to a remote server.

[1564] "High-quality output" refers to highly accurate information and results generated by a large-scale language model that are tailored to the user's needs.

[1565] "Receiving" refers to receiving the generated output from the server at the user's terminal.

[1566] "Displaying" refers to visually presenting the received high quality output to a user.

[1567] "Emotion recognition" refers to the process of determining a user's emotional state from their input.

[1568] "Adjusting data preprocessing" refers to changing and optimizing the content and methods of preprocessing based on the recognized emotions.

[1569] The system for implementing this invention analyzes user input, recognizes emotions, and provides high-quality output based on those emotions. The system consists of two main components: a terminal and a server.

[1570] Program Overview

[1571] 1. Capturing user input

[1572] The user inputs data through a terminal (e.g., a smartphone). The input data can be text data or voice data.

[1573] 2. Emotional Recognition

[1574] The device passes the input data to an emotion engine (e.g., Hugging Face sentiment-analysis) to recognize the user's emotion. For example, if the user input is "Are there any deals?", the device recognizes the user's emotion as "happy" based on that input.

[1575] 3. Pretreatment

[1576] It uses a small language model (e.g., T5-small) to analyze user input and convert it into a concrete form while taking into account emotional information, for example, converting the vague expression "great deal" into "today's sale items," and adjusts the processing speed based on the emotion.

[1577] 4. Data transmission

[1578] The pre-processed data is transmitted to the server by the data transmission module of the terminal.

[1579] 5. Processing with large-scale language models

[1580] The server uses large-scale language models to analyze the pre-processed data and generate high-quality output, such as a list of today's sale items.

[1581] 6. Send and display high-quality output

[1582] The server generates high-quality output and sends it to the terminal, which visually displays the received data to the user.

[1583] Hardware and software used

[1584] Device: User device such as a smartphone or tablet

[1585] Emotion Engine: A sentiment-analysis model of Hugging Face

[1586] Small language models: Natural language processing models such as T5-small

[1587] Server: High-performance server on the cloud

[1588] Large-scale language model: High-performance natural language processing model running on a server

[1589] Specific examples

[1590] The specific scenario is as follows:

[1591] 1. The user speaks into the smartphone app and says, "Do you have any deals?"

[1592] 2. The device captures the input and recognizes the "happy" emotion using its emotion engine.

[1593] 3. Small LLMs concretize the expression "good deal" and transform it into "What's on sale today?"

[1594] 4. Send it to the server, and the large-scale LLM generates a "list of today's sale items" that meets your request.

[1595] 5. Once the list is returned, the device displays it to the user along with "Today's Sale Items."

[1596] Example prompt sentence:

[1597] Q: Are there any great deals? I'd be happy to provide information at a speed that matches my needs. A:

[1598] In this way, the system can provide personalized information based on user sentiment in real time, improving the customer experience.

[1599] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1600] Step 1:

[1601] The user types "Do you have any deals?" into the smartphone app. The input can be text or voice. The device captures this data and stores it as input data.

[1602] Step 2:

[1603] The device passes the captured input data to the emotion engine, which (Hugging Face's sentiment-analysis model) recognizes the emotion from the user's input. In this step, based on the input data (e.g., "Are there any deals available?"), the emotion engine outputs the emotion information "happy."

[1604] Step 3:

[1605] The device takes the recognized emotional information into consideration and passes the user's input to a small-scale LLM (T5-small model). The small-scale LLM analyzes the input and performs data preprocessing based on the emotional information. In this step, the device outputs preprocessed data (e.g., "What are the sale items today?") based on the input data and emotional information (e.g., "Are there any deals?" and "I'm happy").

[1606] Step 4:

[1607] The data transmission module of the device sends the preprocessed data to a server on the cloud, where the preprocessed data ("What are the sale items today?") is transferred to the server as input data.

[1608] Step 5:

[1609] The server passes the received preprocessed data to the large-scale LLM, which analyzes the preprocessed data and generates the corresponding high-quality output. In this step, the server outputs a high-quality output ("List of sale items") based on the preprocessed input data ("What items are on sale today?").

[1610] Step 6:

[1611] The server's data transmission module transmits the generated high-quality output to the terminal, where the generated data ("list of sale items") is transmitted to the terminal.

[1612] Step 7:

[1613] The terminal uses a display module to display the received high-quality data to the user. In this step, the terminal visually presents the high-quality output data ("list of sale items") to the user.

[1614] The above is the specific process flow from user input to the provision of high-quality information. Data is input and output at each step, and the next step is executed based on the results, thereby realizing the provision of highly accurate, emotion-responsive information.

[1615] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1616] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1617] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1618] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1619] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1620] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1621] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1622] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1623] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1624] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1625] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1626] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1627] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1628] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1629] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1630] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1631] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1632] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1633] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1634] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1635] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1636] The following is further disclosed regarding the above embodiment.

[1637] (Claim 1)

[1638] means for parsing and preprocessing user input using a small language model;

[1639] a means for sending the preprocessed data to a large-scale language model on the cloud; and

[1640] means for receiving high quality output generated by the large-scale language model;

[1641] The system includes means for displaying the received high quality output to a user.

[1642] (Claim 2)

[1643] 2. The system according to claim 1, further comprising means for masking and protecting user privacy information in the preprocessing.

[1644] (Claim 3)

[1645] 2. The system according to claim 1, further comprising means for converting questions into a suitable format in the preprocessing step to prevent unnecessary repetition of questions.

[1646] (Claim 4)

[1647] The system according to claim 1, further comprising means for installing a small-scale language model for preprocessing on a terminal and a large-scale language model in a cloud environment.

[1648] (Claim 5)

[1649] 10. The system of claim 1, further comprising means for aggressively leveraging small language models to condition user input into a proper format and improve energy efficiency.

[1650] "Example 1"

[1651] (Claim 1)

[1652] means for analyzing and preprocessing user input using a small language model installed on the user device;

[1653] a means for sending the preprocessed data to a large-scale language model on the cloud; and

[1654] means for receiving high quality output generated by the large-scale language model;

[1655] The system includes means for displaying the received high quality output to a user.

[1656] (Claim 2)

[1657] 2. The system according to claim 1, further comprising a means for anonymizing user privacy information in the preprocessing.

[1658] (Claim 3)

[1659] 10. The system of claim 1, further comprising means for concretizing ambiguous expressions of questions in the preprocessing step.

[1660] (Claim 4)

[1661] 10. The system of claim 1, including means for the large-scale language model to utilize databases and external APIs to gather information and generate high-quality output.

[1662] (Claim 5)

[1663] 10. The system of claim 1, further comprising means for transmitting and receiving the pre-processed data and the high-quality output using a secure communication protocol.

[1664] "Application Example 1"

[1665] (Claim 1)

[1666] means for parsing and preprocessing user input using a small language model;

[1667] a means for sending the preprocessed data to a large-scale language model on the cloud; and

[1668] means for receiving high quality output generated by the large-scale language model;

[1669] means for displaying the received high quality output to a user;

[1670] A means of converting vague expressions into concrete conditions in preprocessing and anonymizing location information;

[1671] A means for masking privacy information to a city / ward / town / village level to protect user privacy;

[1672] a means for generating a list of restaurants in order of highest rating;

[1673] A means of collecting and organizing information based on preprocessed data;

[1674] A means of generating high-quality output from prompts using a generative AI model; and

[1675] A system including:

[1676] (Claim 2)

[1677] 2. The system according to claim 1, further comprising means for masking and protecting user privacy information in the preprocessing.

[1678] (Claim 3)

[1679] 2. The system according to claim 1, further comprising means for converting questions into a suitable format in the preprocessing step to prevent unnecessary repetition of questions.

[1680] "Example 2: Combining Emotion Engines"

[1681] (Claim 1)

[1682] means for parsing and preprocessing user input using a small language model;

[1683] a means for sending the preprocessed data to a large-scale language model on the cloud; and

[1684] means for receiving high quality output generated by the large-scale language model;

[1685] means for displaying the received high quality output to a user;

[1686] a means for capturing user input;

[1687] a means for performing sentiment analysis;

[1688] means for specifying and, if necessary, prioritizing ambiguous expressions based on user input;

[1689] A system including:

[1690] (Claim 2)

[1691] 2. The system according to claim 1, further comprising means for masking and protecting user privacy information in the preprocessing.

[1692] (Claim 3)

[1693] 2. The system according to claim 1, further comprising means for converting questions into a suitable format in the preprocessing step to prevent unnecessary repetition of questions.

[1694] "Application example 2 when combining emotion engines"

[1695] (Claim 1)

[1696] means for parsing and preprocessing user input using a small language model;

[1697] a means for sending the preprocessed data to a large-scale language model on the cloud; and

[1698] means for receiving high quality output generated by the large-scale language model;

[1699] means for displaying the received high quality output to a user;

[1700] means for recognizing emotions from user input;

[1701] a means for adjusting data preprocessing based on the recognized emotion;

[1702] A system including:

[1703] (Claim 2)

[1704] 2. The system according to claim 1, further comprising means for masking and protecting user privacy information in the preprocessing.

[1705] (Claim 3)

[1706] 2. The system according to claim 1, further comprising means for converting questions into a suitable format in the preprocessing step to prevent unnecessary repetition of questions. [Explanation of symbols]

[1707] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for parsing and preprocessing user input using a small language model; a means for sending the preprocessed data to a large-scale language model on the cloud; and means for receiving high quality output generated by the large-scale language model; The system includes means for displaying the received high quality output to a user.

2. 2. The system according to claim 1, further comprising means for masking and protecting user privacy information in the preprocessing.

3. 2. The system according to claim 1, further comprising means for converting queries into a suitable format in the preprocessing step to prevent unnecessary repetition of queries.

4. The system according to claim 1, further comprising means for installing a small-scale language model for preprocessing on a terminal and a large-scale language model in a cloud environment.

5. 10. The system of claim 1, further comprising means for aggressively utilizing small language models to condition user input into a proper format and improve energy efficiency.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A