system

A system using historical data to pre-train and fine-tune a natural language processing model for document generation addresses the inefficiencies of conventional methods, enhancing accuracy and consistency in document creation.

JP2026062120APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional document creation processes are time-consuming, labor-intensive, and lack accuracy and consistency, especially in creating internal company petitions where past data is not effectively utilized.

Method used

A system that collects historical document data to pre-train a natural language processing model, fine-tunes it with specific parameters, and generates document drafts based on user requests, ensuring accuracy and consistency.

Benefits of technology

Significantly reduces time and effort required for document creation while maintaining high accuracy and consistency, allowing efficient document generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062120000001_ABST
    Figure 2026062120000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for collecting past document data, A means for pre-training a natural language processing model using the aforementioned past document data, A means for fine-tuning the natural language processing model using specific parameters, A means for receiving a request from a user and generating a draft document using the finely tuned natural language processing model, Means for providing the generated draft document to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventional document creation processes are often performed manually, which has the problems of taking time and effort. Also, especially in the creation of internal company petitions, since past documents are often referred to, there are cases where past data cannot be effectively utilized. Furthermore, it is also difficult to maintain the accuracy and consistency of document content. Against this background, there is an increasing need for a system that supports efficient and accurate document creation.

Means for Solving the Problems

[0005] This invention provides a means for collecting historical document data and using that data to pre-train a natural language processing model. By fine-tuning this pre-trained model using specific parameters (e.g., monetary amounts and time periods), its ability to generate more specific and accurate document drafts is improved. When a user submits a request, the system generates a document draft using this fine-tuned model and provides it to the user. Furthermore, since the historical document data is extracted from the company's internal database, it is possible to generate document drafts adapted to the company's specific document content and format. In this way, the invention provides a system that significantly reduces time and effort while guaranteeing the accuracy and consistency of document content.

[0006] "Past document data" refers to information about documents created in the past, extracted from specific databases or similar sources.

[0007] A "natural language processing model" is a type of machine learning model designed to understand and generate human language.

[0008] "Pre-training" refers to the process of training a model with a large amount of data to learn common patterns and vocabulary.

[0009] "Fine-tuning learning" is a process of further training a pre-trained model using specific parameters to improve the model's accuracy and adaptability.

[0010] "Specific parameters" refer to the specific conditions or numerical settings that the system uses to generate documents, and include, for example, monetary amounts and time periods.

[0011] "User requests" refer to the demands or inputs that users make to the system.

[0012] "Draft document" refers to automatically generated drafts or template documents.

[0013] "Decoding" refers to the process of converting tokenized text data generated by a machine learning model into a human-readable text format.

[0014] A "database" refers to a system for managing, searching, and using stored information.

[0015] A "proposal document" refers to an official document submitted within a company or organization to obtain approval for a specific matter. [Brief explanation of the drawing]

[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0020] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] The system of the present invention automatically generates draft documents based on user requests by collecting historical document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. This system is implemented according to the following procedure.

[0038] Program processing

[0039] 1. The server first collects historical document data from the company's internal database. This includes various official documents, such as approval forms. The collected data is saved in text file or CSV format.

[0040] 2. The server then sets up a natural language processing model (e.g., BERT). It loads the collected historical document data and uses this data to pre-train the model. It tokenizes the document data using a tokenizer and feeds it to the model as input to learn patterns and vocabulary.

[0041] 3. The server performs fine-tuning learning on the pre-trained model using specific parameters (e.g., amount or timing). This allows the model to adapt to a specific business context and generate more specific and accurate document drafts. Fine-tuning learning is also performed by tokenizing data using a tokenizer and feeding it to the model.

[0042] 4. The user sends a request to the chatbot to create an approval document using their own device. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0043] 5. The terminal sends the user's request to the server. The server receives this request, tokenizes the request content, and provides it to the refined model.

[0044] 6. The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Specifically, the content generated by the model is processed into a draft document and converted into a format that the user can review.

[0045] 7. The terminal displays the generated document draft to the user. This allows the user to obtain a suitable document draft in a short amount of time.

[0046] Specific example

[0047] For example, if a user sends a request to the chatbot to create an approval document for a payment of 5 million yen on December 15, 2023, the following process will occur:

[0048] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0049] 2. The device sends this request to the server.

[0050] 3. The server tokenizes the request and feeds it to a finely tuned natural language processing model.

[0051] 4. The server generates a draft approval document and decodes it as follows: "Proposal: The budget for this project is 5 million yen, and the implementation date is December 15, 2023."

[0052] 5. The terminal displays this draft document to the user.

[0053] In this way, the system of the present invention provides an environment in which users can easily create approval documents, thereby reducing time and effort.

[0054] The following describes the processing flow.

[0055] Step 1:

[0056] The server collects historical document data from the company's internal database. This includes data from official documents such as approval requests and reports. The collected data is stored in a single text file or a CSV file.

[0057] Step 2:

[0058] The server initializes a natural language processing model (e.g., BERT). This involves using a tool called a tokenizer to tokenize the collected text data (divide it into units of words or sentences) and convert it into a format that the model can understand.

[0059] Step 3:

[0060] The server inputs tokenized data into a natural language processing model for pre-training. This allows the model to learn common document patterns and vocabulary from past document data.

[0061] Step 4:

[0062] The server prepares specific parameters (e.g., amount and time period). Based on these specific parameters, it further refines the pre-trained model. This improves the model's accuracy and adaptability to the specific business context.

[0063] Step 5:

[0064] The user sends a request to the chatbot from their device to create an approval document. This request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0065] Step 6:

[0066] The terminal sends data to the server via a communication interface in order to send user requests to the server.

[0067] Step 7:

[0068] The server tokenizes the user's request using a tokenizer. This tokenized data is then input into a finely tuned natural language processing model.

[0069] Step 8:

[0070] The server decodes the tokenized data generated from the model and creates a draft document in an easy-to-read format. Specifically, it converts the generated tokenized data into natural language sentences.

[0071] Step 9:

[0072] The terminal displays the draft document sent from the server to the user. This allows the user to quickly review an appropriate draft document.

[0073] This series of processes allows users to efficiently create approval documents, significantly reducing the time and effort required.

[0074] (Example 1)

[0075] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0076] In modern businesses, document creation is a time-consuming and labor-intensive task. In particular, creating new documents based on past documents is cumbersome, requiring careful reference and editing, and maintaining consistency is difficult. Furthermore, generating draft documents quickly and accurately in response to specific user requests is challenging to do manually. This leads to a decrease in overall business efficiency.

[0077] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0078] In this invention, the server includes means for collecting historical document data, means for pre-training a natural language processing model using the historical document data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating document drafts using the fine-tuning natural language processing model, means for providing the generated document drafts to users, means for tokenizing the content of the user's request, means for inputting the tokenized data into the natural language processing model, and means for decoding the tokenized data generated from the natural language processing model. This makes it possible for users to easily and efficiently create high-quality documents.

[0079] "Past document data" refers to information about previously created documents extracted from a company's internal document database.

[0080] A "natural language processing model" is a machine learning model that analyzes and understands text data, and in particular, it uses deep learning techniques to learn the meaning of sentences.

[0081] "Pre-training" is the process of training a natural language processing model using a large amount of text data to learn basic language patterns and structures.

[0082] "Fine-tuning learning" is the process of setting specific parameters for a pre-trained natural language processing model and making adjustments to suit a more specific context.

[0083] "Specific parameters" refer to values ​​or conditions (e.g., amount, time) that are important elements when generating a document.

[0084] A "user request" refers to instructions or requests regarding document generation that a user sends to the system via their device.

[0085] "Tokenization" is the process of dividing text data into smaller units (tokens) such as words and phrases.

[0086] "Decoding" is the process of converting tokenized data back into natural language text in a format that is easy for humans to read.

[0087] A "draft document" is a preliminary version of a document generated by a natural language processing model and created based on user requests.

[0088] This invention is a system for efficiently creating documents within a company. It collects past document data, pre-trains a natural language processing model using that data, and then performs fine-tuning learning using specific parameters, thereby automatically generating document drafts based on user requests.

[0089] First, the server collects historical document data from the company's internal database. This data includes various types of documents such as approval forms, reports, and contracts. The collected data is saved in text or CSV format in preparation for the next processing step.

[0090] Next, the server prepares a natural language processing model (e.g., BERT) using the Python Torch library. It loads historical document data and uses this data to pre-train the model. Specifically, it tokenizes the document data using a tokenizer (e.g., BertTokenizer) and inputs this tokenized data into the model. This allows the model to learn document patterns and vocabulary.

[0091] Once pre-training is complete, the server fine-tunes the model using specific parameters (e.g., amount or time period). This fine-tuning is also performed by tokenizing data using a tokenizer and feeding it to the model. This allows the model to generate more specific and accurate document drafts adapted to the particular business context.

[0092] The user sends a request to the chatbot to create an approval document using their device. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023"). The device then sends this request to the server.

[0093] The server tokenizes the received request and feeds it to a finely tuned model. The model generates a draft document based on the request. The generated tokenized data is decoded again into a readable draft document. Finally, the terminal displays this draft document to the user.

[0094] Specific example

[0095] For example, if a user sends a request to the chatbot to create an approval document for a payment of 5 million yen on December 15, 2023, the following process will occur:

[0096] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0097] 2. The device sends this request to the server.

[0098] 3. The server tokenizes the request and feeds it to a finely tuned natural language processing model.

[0099] 4. The server generates a draft approval document using a natural language processing model.

[0100] 5. The server decodes the generated content and creates a draft document stating, "Proposed content: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023."

[0101] 6. The terminal displays this draft document to the user.

[0102] This system allows users to obtain appropriate and consistent document drafts in a short amount of time, thereby improving work efficiency.

[0103] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0104] Step 1:

[0105] The server executes a mechanism to collect historical document data from the company's internal database. It uses database queries to retrieve document data in text or CSV format, such as approval forms, reports, and contracts. The database queries used as input output document data, which is saved to the "past_documents.csv" file. Specifically, it calls an API for database access and executes SQL queries.

[0106] Step 2:

[0107] The server prepares and configures a natural language processing model (e.g., BERT). It loads the model using the Python Torch library and sets the tokenizer, for example, `tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')`. It reads a past document data file, "past_documents.csv", as input and tokenizes it using the tokenizer. The output is the tokenized document data.

[0108] Step 3:

[0109] The server pre-trains a BERT model using past document data. The tokenized data from the tokenizer is created as a dataset using `dataset = TextDataset(tokenizer, 'past_documents.csv')`, and the model is pre-trained using `model.train(dataset)`. The pre-trained model is then output using the tokenized data output from the tokenizer as input.

[0110] Step 4:

[0111] The server performs fine-tuning on a pre-trained model using specific parameters (e.g., amount or time period). It sets specific parameters, tokenizes the fine-tuning data using a tokenizer, and feeds it to the model. Specifically, it executes `model.fine_tune(tokenized_data, parameters)`. The input is the specific parameters and tokenized fine-tuning data, and the output is the fine-tuned model.

[0112] Step 5:

[0113] The user sends a request for the creation of an approval document to the chatbot via their device. Specifically, they enter "Amount: 5 million yen, Date: December 15, 2023" into the chatbot's interface and press the send button. The input is the user's prompt text, and the output is the request sent by the chatbot's interface.

[0114] Step 6:

[0115] The terminal sends the user's request to the server. Specifically, it creates an HTTP POST request and sends the request data to a specific endpoint on the server (e.g., " / generate_document"). The input is the user's request, and the output is the sending of the HTTP request to the server.

[0116] Step 7:

[0117] The server tokenizes the request content and provides it to a fine-tuned model. The received request is tokenized using `tokenizer.tokenize(request_data)`, and a draft document is generated using `model.generate(tokenized_request)`. The input is the tokenized request data, and the output is the generated draft document.

[0118] Step 8:

[0119] The server decodes the generated tokenized data and creates a draft document. The token output is decoded using tokenizer.decode(generated_tokens) and converted into a user-readable format. The input is the tokenized draft document data, and the output is the decoded draft document.

[0120] Step 9:

[0121] The terminal displays the generated document draft to the user. Specifically, it receives the document draft from the server and displays it on the chatbot screen. The input is the document draft from the server, and the output is the display to the user.

[0122] (Application Example 1)

[0123] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0124] Logistics centers require a wide variety of documents (purchase orders, shipping instructions, inventory management reports, etc.), but creating these documents is time-consuming and labor-intensive. Furthermore, because these documents must accurately reflect past data and current conditions, they are prone to human error. While efficiency is needed in these document creation processes, on-site personnel often require advanced specialized knowledge. A system is needed to address these challenges.

[0125] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0126] In this invention, the server includes means for collecting historical data, means for pre-training a natural language processing model using the historical data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating draft documents using the fine-tuning natural language processing model, means for providing the generated draft documents to users, means for collecting historical order history and inventory data from the logistics center's database, means for automatically generating purchase orders, shipping instructions, and inventory management reports using the natural language processing model, and means for displaying and modifying the generated draft documents on a smartphone. This makes it possible to automatically generate documents efficiently and accurately at the logistics center, significantly reducing the burden on on-site staff and improving work efficiency.

[0127] "Past data" refers to information including past order history and inventory data at the logistics center.

[0128] A "natural language processing model" is a machine learning model that understands human language and generates documents.

[0129] "Pre-training" is the process of training a natural language processing model from the ground up using a large amount of data.

[0130] "Fine-tuning learning" is the process of optimizing a natural language processing model, which has been pre-trained using specific parameters, to suit a particular application.

[0131] A "user request" is information that a user uses to request document generation based on specific parameters.

[0132] A "draft document" is a document generated by a natural language processing model and submitted as a candidate.

[0133] "Means of provision" refers to the method by which the generated draft document is displayed for the user to review and modify.

[0134] A "logistics center database" is a data system that stores order history and inventory data managed within a logistics center.

[0135] A "purchase order" is a document that details the order for goods or services.

[0136] A "shipping instruction sheet" is a document that contains instructions for shipping goods out of a logistics center.

[0137] An "inventory management report" is a report that summarizes the current inventory status and related information.

[0138] A "smartphone" is a portable, multi-functional device that can display and process information using applications.

[0139] This invention is a system for efficiently and automatically generating documents in a logistics center. This system collects historical data, uses that data to pre-train a natural language processing model, and then fine-tunes it using specific parameters. Next, it generates draft documents based on user requests and provides them to the user.

[0140] The specific configuration of this system is as follows:

[0141] Hardware and software:

[0142] Server: Responsible for data collection, model training, fine-tuning, and document generation. The natural language processing model used includes BERT.

[0143] Database system: Stores order history and inventory data for the logistics center. Uses MySQL® or PostgreSQL, etc.

[0144] Tokenizer: Uses a Python library (such as Hugging Face's Transformers) to tokenize the collected data.

[0145] User interface device (smartphone): Built with React Native to display generated document drafts and receive user requests.

[0146] Program processing:

[0147] 1. Data collection:

[0148] The server collects past order history and inventory data from the logistics center's database. The data is saved in text file or CSV format.

[0149] 2. Pre-learning:

[0150] The server sets up a natural language processing model (BERT) and pre-trains the model using collected historical data. The data is tokenized using the Python Hugging Face library and input into the model.

[0151] 3. Fine-tuning learning:

[0152] The server fine-tunes a pre-trained model using specific parameters (product name, quantity, delivery date). This process also uses a tokenizer to tokenize the data and add optimization information to the model.

[0153] 4. Receiving and sending user requests:

[0154] Users use their smartphones to submit requests for new purchase orders or shipping orders. These requests include the necessary parameters (e.g., "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023").

[0155] The smartphone sends this request to the server.

[0156] 5. Document draft generation:

[0157] The server tokenizes the received request and feeds it into a finely tuned model. The model generates a draft document based on the specified parameters.

[0158] 6. Providing draft documents:

[0159] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. For example, it might generate a specific draft document such as, "Please deliver 100 units of product A on December 1, 2023."

[0160] The generated document draft is displayed to the user on their smartphone. The user can review the content and make corrections as needed.

[0161] This will enable the automation and streamlining of document generation tasks in logistics centers. Furthermore, it will allow users to obtain accurate documents in a shorter time, improving both work efficiency and accuracy.

[0162] Specific example:

[0163] When a user submits a request to create a purchase order with the following information: "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023," the server tokenizes this information and feeds it into a finely tuned natural language processing model to generate a purchase order like this: "Please deliver 100 units of Product A on December 1, 2023." This draft document is then displayed to the user on their smartphone.

[0164] Example of a prompt:

[0165] "Please create a purchase order using the following parameters: Product Name: Product A, Quantity: 100, Due Date: December 1, 2023"

[0166] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0167] Step 1:

[0168] The server collects past order history and inventory data from the logistics center's database. It queries the database system (e.g., MySQL or PostgreSQL) to extract the necessary data. The collected data is saved in text file or CSV format.

[0169] Input: Logistics center database

[0170] Output: Text or CSV file containing past order history and inventory data.

[0171] Step 2:

[0172] The server pre-trains a natural language processing model (BERT) using the collected historical data. The data is tokenized using the Python Hugging Face library and input into the model. This allows the model to learn basic language structures and patterns.

[0173] Input: Text or CSV file containing past order history and inventory data.

[0174] Output: Pre-trained natural language processing model

[0175] Step 3:

[0176] The server fine-tunes a pre-trained model using specific parameters (product name, quantity, delivery date). The Python Hugging Face library is used to tokenize additional data and add optimization information to the model. This allows the model to adapt to specific business contexts.

[0177] Input: Specific parameters (product name, quantity, delivery date)

[0178] Output: Fine-tuned natural language processing model

[0179] Step 4:

[0180] Users use their smartphones to submit requests for new purchase orders or shipping orders. These requests include the necessary parameters (e.g., "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023").

[0181] Input: Request parameters (product name, quantity, delivery date)

[0182] Output: Request from the user's smartphone

[0183] Step 5:

[0184] The terminal sends the user's request to the server. The request data is transferred to the server using protocols such as HTTP requests.

[0185] Input: Request from the user's smartphone

[0186] Output: Request data to the server

[0187] Step 6:

[0188] The server tokenizes the received request using a tokenizer and feeds it into a finely tuned model. The model generates a draft document based on the specified parameters.

[0189] Input: Request data to the server

[0190] Output: Tokenized data of the generated document

[0191] Step 7:

[0192] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Specifically, it converts the tokenized data into natural language sentences and formats it as a structured document.

[0193] Input: Tokenized data of the generated document

[0194] Output: Draft document in an easy-to-read format

[0195] Step 8:

[0196] The terminal displays the generated document draft to the user. The user can review the document draft through the smartphone interface and make revisions as needed.

[0197] Input: Draft document in an easy-to-read format

[0198] Output: Draft document that users can review and edit.

[0199] The above steps automate and streamline document generation at the logistics center. Clearly defining the specific actions and data input / output performed at each step makes system implementation and operation easier.

[0200] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0201] The system of the present invention automatically generates draft documents based on user requests by collecting past document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the content and tone of the draft documents can be adjusted according to the user's emotions. This system is implemented according to the following procedure.

[0202] Program processing

[0203] 1. The server collects historical document data from the company's internal database. The collected data is saved in text file or CSV format.

[0204] 2. The server sets up a natural language processing model (e.g., BERT). The collected document data is tokenized using a tokenizer and input into the model. The model is then pre-trained to learn common document patterns and vocabulary.

[0205] 3. The server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning allows the model to adapt to the specific business context and improves its accuracy.

[0206] 4. The user sends a request to the chatbot from their device to create a proposal document. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0207] 5. The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes emotions from the input text or voice and sends the results to the server.

[0208] 6. The server tokenizes the user's request based on the sentiment analysis results received from the sentiment engine and inputs it into a finely tuned model. The model generates a draft document taking the sentiment analysis results into account.

[0209] 7. The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Based on the sentiment analysis results, the tone and content of the document are adjusted.

[0210] 8. The terminal displays the generated draft document to the user. This allows the user to quickly obtain a draft document that takes into account an appropriate and emotionally balanced tone.

[0211] Specific example

[0212] For example, if a user submits a request to create a proposal document for a transaction with the following conditions: "Amount: 5 million yen, Date: December 15, 2023, Emotion: Angry," the following process will occur:

[0213] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0214] 2. The device sends this request to the emotion engine to analyze the user's emotions. If the analysis result is "anger," the tone is taken into consideration.

[0215] 3. The device sends the request to the server along with the sentiment analysis results.

[0216] 4. The server tokenizes the request and feeds it into a finely tuned natural language processing model. The model generates a draft document, taking sentiment analysis results into account.

[0217] 5. The server decodes the draft document and generates content such as, "Proposal: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[0218] 6. The terminal displays this draft document to the user.

[0219] This series of processes allows users to efficiently create approval documents that take emotions into consideration, thereby improving the accuracy and effectiveness of communication.

[0220] The following describes the processing flow.

[0221] Step 1:

[0222] The server collects historical document data from the company's internal database. The collected data is saved in text file or CSV format.

[0223] Step 2:

[0224] The server configures a natural language processing model (e.g., BERT). It tokenizes the collected document data using a tokenizer. The tokenizer divides the document data into units of words and phrases, converting it into a format that the model can understand.

[0225] Step 3:

[0226] The server inputs tokenized data into a natural language processing model for pre-training. The model learns common document patterns and vocabulary from past document data.

[0227] Step 4:

[0228] The server performs fine-tuning learning on a pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning allows the model to adapt to a specific business context, improving its accuracy and adaptability.

[0229] Step 5:

[0230] The user sends a request to the chatbot from their device to create an approval document. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0231] Step 6:

[0232] The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the emotions from the input text or voice and sends the results to the server.

[0233] Step 7:

[0234] The server tokenizes the user's request based on the sentiment analysis results received from the sentiment engine and inputs it into a finely tuned model. The model then generates a draft document, taking the sentiment analysis results into account.

[0235] Step 8:

[0236] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. The generated draft document is then adjusted in tone and content based on the sentiment analysis results.

[0237] Step 9:

[0238] The terminal displays the generated draft document to the user. This allows the user to quickly review a draft document that takes into account appropriate and emotional tone.

[0239] Specific example

[0240] For example, if a user submits a request to create a proposal document for a transaction with the following conditions: "Amount: 5 million yen, Date: December 15, 2023, Emotion: Angry," the following process will occur:

[0241] Step 1:

[0242] The user submits a request from their device to create an approval document for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0243] Step 2:

[0244] The device sends this request to the emotion engine, which analyzes the user's emotions. The analysis results detect that the user is "angry."

[0245] Step 3:

[0246] The emotion engine sends the emotion analysis results to the server.

[0247] Step 4:

[0248] The server tokenizes the request along with the sentiment analysis results and inputs them into a finely tuned model. The model then generates a draft document, taking the sentiment analysis results into account.

[0249] Step 5:

[0250] The server decodes the generated draft document and produces content such as, "Proposal: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[0251] Step 6:

[0252] The terminal displays the generated draft document to the user.

[0253] In this way, users can create emotionally conscious document drafts quickly and efficiently.

[0254] (Example 2)

[0255] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0256] Conventional document generation systems failed to consider user emotions, resulting in a uniform tone in the generated documents. This led to the problem of documents not adequately reflecting user intent and feelings. Furthermore, methods for providing highly accurate models adapted to business contexts through fine-tuning learning using past document data were insufficient.

[0257] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0258] In this invention, the server includes means for collecting past document data, means for pre-training a natural language processing model using the past document data, means for fine-tuning the natural language processing model using specific parameters, means for analyzing user requests to recognize emotional states, and means for adjusting the tone and content of draft documents considering the emotional states. This makes it possible to generate documents with an appropriate tone that reflects the user's emotions. Furthermore, it is possible to provide highly accurate documents adapted to a specific business context.

[0259] "Past document data" refers to documents previously created by companies or individuals that are stored in databases or file systems.

[0260] A "natural language processing model" is an artificial intelligence model used to understand and generate human language, and includes models such as BERT.

[0261] "Pre-training" is the process of initially training a model using a large amount of collected data to teach it common language patterns and vocabulary.

[0262] "Fine-tuning learning" is the process of further training a pre-trained model according to specific tasks or contexts to improve the model's accuracy.

[0263] A "user request" is input from a user to give specific instructions or requests to the system, and is sent through a chatbot or form.

[0264] A "document draft" is a draft of a document, such as a proposal or report, generated through user requests and system processing.

[0265] "Emotional state" refers to the state of emotions analyzed from the user's input, and includes emotions such as anger, joy, and sadness.

[0266] "Tokenization" is the process of breaking down text data into units of words or subwords and assigning an ID to each of them.

[0267] An "emotion engine" is software or an algorithm that analyzes a user's text or voice input to recognize their emotional state.

[0268] "Decoding" is the process of converting tokenized data back into a human-readable text format.

[0269] The system of this invention automatically generates draft documents based on user requests by collecting past document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the content and tone of the draft documents can be adjusted according to the user's emotions.

[0270] Specifically, the server first collects historical document data from the company's internal database. This data includes business documents such as sales reports, meeting minutes, and approval documents. The collected data is saved in text file or CSV format.

[0271] Next, the server loads a natural language processing model (e.g., BERT) using the Hugging Face library. The collected document data is tokenized using a tokenizer (e.g., BERT Tokenizer) and input into the model. The model is pre-trained using this tokenized data to learn common language patterns and vocabulary.

[0272] Subsequently, the server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount and timing). This fine-tuning allows the model to adapt to the specific business context, enabling more accurate document generation. The fine-tuned model is stored on the server and used to respond to user requests.

[0273] The user uses a chatbot interface from their device to request the creation of an approval document for specific details (e.g., "Amount: 5 million yen, Date: December 15, 2023"). This request is sent from the device to the emotion engine, which analyzes the user's emotional state based on their input. The emotion engine classifies the user's emotions into categories such as "anger," "joy," and "sadness," and returns the result to the device.

[0274] The terminal sends the sentiment analysis results and the user's request to the server, which tokenizes the request and inputs it into a finely tuned model. The model takes the sentiment analysis results into account and generates a draft document that reflects the user's emotions. For example, if anger is detected, the tone of the document is strengthened and adjusted to convey a request for immediate action.

[0275] The generated tokenized data is decoded on the server and sent to the terminal as a readable, formatted draft document. The terminal displays the final draft document to the user, who then reviews the generated document and makes any necessary corrections or additions.

[0276] As a concrete example, consider a case where a user submits a request to create a proposal document with the following conditions: "Payment amount: 5 million yen, Date: December 15, 2023, Emotion: Angry." In response to this request, the server generates a draft document reflecting the user's emotion and provides it to the user. The generated draft document would be displayed with content such as: "Proposal: The payment amount for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[0277] This system allows users to efficiently create documents that take emotions into account, thereby improving the accuracy and effectiveness of communication.

[0278] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0279] Step 1:

[0280] The server collects historical document data from the company's internal database. This includes business documents such as sales reports, meeting minutes, and approval documents. The collected data is saved in text file or CSV format.

[0281] Input: Corporate database

[0282] Output: Collected historical document data (text files, CSV)

[0283] Step 2:

[0284] The server loads a natural language processing model (e.g., BERT) using the Hugging Face library. It tokenizes the collected document data using a tokenizer (e.g., BERT Tokenizer). Each document is decomposed into words or sub-word units and converted into integer-valued IDs.

[0285] Input: Past document data

[0286] Output: Tokenized document data

[0287] Step 3:

[0288] The server inputs the tokenized data into the model for pre-training. In this process, the model learns the patterns and vocabulary of the documents. An initially trained model is obtained.

[0289] Input: Tokenized document data

[0290] Output: Pre-trained natural language processing model

[0291] Step 4:

[0292] The server performs fine-tuning on the pre-trained model using specific parameters (e.g., amount, time). As a result, the model adapts to a specific business context and its accuracy is improved.

[0293] Input: Pre-trained natural language processing model, specific parameters

[0294] Output: Fine-tuned natural language processing model

[0295] Step 5: <​​​​

[0297] Input: User request (such as payment amount, time, etc.)

[0298] Output: Request data received by the terminal

[0299] Step 6:

[0300] The terminal sends the user's request to the emotion engine to analyze the user's emotion. The emotion engine classifies the emotional state from the input text or voice and returns the result to the terminal.

[0301] Input: User request, emotion engine

[0302] Output: Emotion analysis result (anger, joy, sadness, etc.)

[0303] Step 7:

[0304] The terminal sends the emotion analysis result and the user's request to the server.

[0305] Input: Emotion analysis result, user request

[0306] Output: Emotion analysis result and request received by the server

[0307] Step 8:

[0308] The server tokenizes the request content and inputs it into the fine-tuned model. The model generates a document draft considering the emotion analysis result and reflecting the user's emotion.

[0309] Input: Tokenized request content, emotion analysis result, fine-tuned natural language processing model

[0310] Output: Generated document draft

[0311] Step 9:<00009​​The server decodes the generated tokenized data and creates a readable, formatted draft document. Based on the sentiment analysis results, the tone and content of the document are adjusted.

[0313] Input: Generated tokenized data

[0314] Output: Formatted draft document

[0315] Step 10:

[0316] The terminal displays the final draft document to the user. The user can review the generated document and make corrections or additions as needed.

[0317] Input: Formatted draft document

[0318] Output: Draft document displayed to the user

[0319] (Application Example 2)

[0320] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0321] Traditional customer service systems have struggled to analyze customer emotions in real time and respond accordingly. Furthermore, relying on manuals makes it difficult to improve customer satisfaction. A system is needed to solve these problems and provide more advanced customer service.

[0322] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0323] In this invention, the server includes means for collecting past document data, means for pre-training a natural language processing model using the past document data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating draft documents using the fine-tuning natural language processing model, means for providing the generated draft documents to the user, and means for analyzing the customer's emotions from their statements and facial expressions, generating a draft document based on the analysis results, and displaying it on a display device. This makes it possible to analyze the customer's emotions in real time and generate and provide an appropriate draft document based on the results.

[0324] "Past document data" refers to existing document information collected from within or outside the company, which is used for pre-training and analysis of the model.

[0325] A "natural language processing model" is a machine learning model designed to enable human understanding of natural language, and is a technology for automatically generating, analyzing, translating, and summarizing document content.

[0326] "Pre-training" is the process of using a broad dataset to teach a model basic language patterns and vocabulary before fine-tuning it for a specific task.

[0327] "Fine-tuning" is an additional learning process that uses specific parameters to further adapt a pre-trained model to a particular task or context.

[0328] A "user request" refers to the specific requests and parameters that a user provides to the system, and is the input information that the system uses to generate appropriate output based on these requests.

[0329] "Customer statements" refer to the words and sentences that customers make to the system, and by analyzing them, we can understand their intentions and requirements.

[0330] "Customer facial expressions" refer to the emotions and reactions shown by a customer's facial expressions, and by analyzing these, information is used to infer the customer's current emotional state.

[0331] "Emotional analysis" refers to the process of automatically identifying customer emotions from text and image data, and then generating appropriate responses and actions based on those results.

[0332] "Generating document drafts" means automatically creating appropriate document content using a natural language processing model based on user requests and analysis results.

[0333] "Displaying on a display device" means displaying information on a device that provides users with a visual representation of the generated document draft or analysis results, such as smart glasses or a display screen.

[0334] The present invention provides a system for pre-training a natural language processing model based on past document data, generating draft documents based on user requests and sentiment analysis results, and displaying them on a display device. The configuration and method for implementing this system are described below.

[0335] System Configuration

[0336] The system of the present invention consists of the following main elements:

[0337] 1. Server

[0338] Data collection method: Collect and store historical document data.

[0339] Natural language processing model: Pre-training is performed using collected document data, and then fine-tuning is performed using specific parameters (e.g., amount, timing).

[0340] Emotion analysis engine: Analyzes customer emotions from their statements and facial expressions.

[0341] 2. Terminal

[0342] Input method: Receives requests from users.

[0343] Display method: The generated document draft is displayed using smart glasses or a display.

[0344] Data collection and learning

[0345] The server collects historical document data from the company's database. This data is stored in text file or CSV format. The collected data is used to pre-train a natural language processing model (e.g., BERT). During pre-training, a tokenizer is used to tokenize the data, which is then input into the model. Next, the pre-trained model is fine-tuned with specific parameters.

[0346] Sentiment analysis and document generation

[0347] The user sends a request through a device (e.g., smart glasses). The request includes parameters such as "Payment amount: 5 million yen, Date: December 15, 2023". The device sends the user's speech and facial expressions to an emotion analysis engine to analyze the customer's emotions. Once the emotion analysis results are obtained, they are sent to the server.

[0348] The server generates a draft document using a finely tuned natural language processing model based on the received sentiment analysis results and request content. During this process, it adjusts the tone and content of the document, taking the sentiment analysis results into consideration. The generated draft document is then sent back to the terminal and displayed on the display device.

[0349] Hardware and software to use

[0350] 1. Hardware:

[0351] Smart glasses: Display customer information and analysis results in real time.

[0352] Webcam: Used to analyze customer facial expressions.

[0353] 2. Software:

[0354] Natural language processing model: BERT is used to analyze user utterances.

[0355] Sentiment analysis engine: Uses the TextBlob library to analyze emotions from customer statements.

[0356] OpenCV: Used to analyze customer facial expressions from image data.

[0357] Specific example

[0358] For example, if a customer says, "I'm having trouble with this product," the following process will occur:

[0359] 1. The user (store clerk) wears smart glasses and interacts with the customer.

[0360] 2. The user's spoken content and facial expression data captured by the webcam are sent to the emotion analysis engine, which then analyzes their emotions.

[0361] 3. The sentiment analysis results and request details are sent to the server, and a draft document is generated.

[0362] 4. The generated draft document is displayed on the smart glasses' screen.

[0363] Examples of prompts to input into a generative AI model include:

[0364] A customer asked a question about our company's products. They seemed confused, so please generate a reassuring message by providing FAQs and relevant product information.

[0365] This method allows users to obtain draft documents in real time that take into account appropriate and emotional tone, enabling them to provide customers with a more satisfying service.

[0366] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0367] Step 1:

[0368] The server collects historical document data from the company's database. This includes data in text file and CSV format. The collected document data is used as input to pre-train a natural language processing model.

[0369] Step 2:

[0370] The server tokenizes the collected document data using a tokenizer and inputs it into a natural language processing model (e.g., BERT). The model is then pre-trained to learn common document patterns and vocabulary. The output is the pre-trained model.

[0371] Step 3:

[0372] The server performs fine-tuning training on a pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning adapts the model to a specific business context, improving its accuracy. The output is the fine-tuned model.

[0373] Step 4:

[0374] The user sends a request from their device. The request includes the necessary parameters (e.g., "Payment amount: 5 million yen, Date: December 15, 2023"). The device sends the user's request to the sentiment analysis engine. The input is the user's request, and the output is the data transfer to the sentiment analysis engine.

[0375] Step 5:

[0376] The device analyzes the user's speech and facial expressions using an emotion analysis engine. The emotion analysis engine analyzes emotions from input text and image data and sends the results to the server. The input is the user's speech and facial expression data, and the output is the emotion analysis result.

[0377] Step 6:

[0378] The server inputs the generated document draft, based on the sentiment analysis results and the user's request, into a finely tuned natural language processing model. The model generates the document draft taking the sentiment analysis results into account. The input is the sentiment analysis results and the user's request, and the output is the generated document draft.

[0379] Step 7:

[0380] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Based on the sentiment analysis results, the tone and content of the document are adjusted. The input is the generated tokenized data, and the output is the draft document in an easy-to-read format.

[0381] Step 8:

[0382] The terminal displays the generated document draft on a display device (e.g., smart glasses or a screen). This allows the user to quickly obtain a document draft that is appropriate and considers emotional tone. The input is a readable document draft, and the output is a display on a display device.

[0383] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0384] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0385] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0386] [Second Embodiment]

[0387] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0388] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0389] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0390] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0391] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0392] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0393] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0394] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0395] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0396] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0397] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0398] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0399] The system of the present invention automatically generates draft documents based on user requests by collecting historical document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. This system is implemented according to the following procedure.

[0400] Program processing

[0401] 1. The server first collects historical document data from the company's internal database. This includes various official documents, such as approval forms. The collected data is saved in text file or CSV format.

[0402] 2. The server then sets up a natural language processing model (e.g., BERT). It loads the collected historical document data and uses this data to pre-train the model. It tokenizes the document data using a tokenizer and feeds it to the model as input to learn patterns and vocabulary.

[0403] 3. The server performs fine-tuning learning on the pre-trained model using specific parameters (e.g., amount or timing). This allows the model to adapt to a specific business context and generate more specific and accurate document drafts. Fine-tuning learning is also performed by tokenizing data using a tokenizer and feeding it to the model.

[0404] 4. The user sends a request to the chatbot to create an approval document using their own device. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0405] 5. The terminal sends the user's request to the server. The server receives this request, tokenizes the request content, and provides it to the refined model.

[0406] 6. The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Specifically, the content generated by the model is processed into a draft document and converted into a format that the user can review.

[0407] 7. The terminal displays the generated document draft to the user. This allows the user to obtain a suitable document draft in a short amount of time.

[0408] Specific example

[0409] For example, if a user sends a request to the chatbot to create an approval document for a payment of 5 million yen on December 15, 2023, the following process will occur:

[0410] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0411] 2. The device sends this request to the server.

[0412] 3. The server tokenizes the request and feeds it to a finely tuned natural language processing model.

[0413] 4. The server generates a draft approval document and decodes it as follows: "Proposal: The budget for this project is 5 million yen, and the implementation date is December 15, 2023."

[0414] 5. The terminal displays this draft document to the user.

[0415] In this way, the system of the present invention provides an environment in which users can easily create approval documents, thereby reducing time and effort.

[0416] The following describes the processing flow.

[0417] Step 1:

[0418] The server collects historical document data from the company's internal database. This includes data from official documents such as approval requests and reports. The collected data is stored in a single text file or a CSV file.

[0419] Step 2:

[0420] The server initializes a natural language processing model (e.g., BERT). This involves using a tool called a tokenizer to tokenize the collected text data (divide it into units of words or sentences) and convert it into a format that the model can understand.

[0421] Step 3:

[0422] The server inputs tokenized data into a natural language processing model for pre-training. This allows the model to learn common document patterns and vocabulary from past document data.

[0423] Step 4:

[0424] The server prepares specific parameters (e.g., amount and time period). Based on these specific parameters, it further refines the pre-trained model. This improves the model's accuracy and adaptability to the specific business context.

[0425] Step 5:

[0426] The user sends a request to the chatbot from their device to create an approval document. This request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0427] Step 6:

[0428] The terminal sends data to the server via a communication interface in order to send user requests to the server.

[0429] Step 7:

[0430] The server tokenizes the user's request using a tokenizer. This tokenized data is then input into a finely tuned natural language processing model.

[0431] Step 8:

[0432] The server decodes the tokenized data generated from the model and creates a draft document in an easy-to-read format. Specifically, it converts the generated tokenized data into natural language sentences.

[0433] Step 9:

[0434] The terminal displays the draft document sent from the server to the user. This allows the user to quickly review an appropriate draft document.

[0435] This series of processes allows users to efficiently create approval documents, significantly reducing the time and effort required.

[0436] (Example 1)

[0437] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0438] In modern businesses, document creation is a time-consuming and labor-intensive task. In particular, creating new documents based on past documents is cumbersome, requiring careful reference and editing, and maintaining consistency is difficult. Furthermore, generating draft documents quickly and accurately in response to specific user requests is challenging to do manually. This leads to a decrease in overall business efficiency.

[0439] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0440] In this invention, the server includes means for collecting historical document data, means for pre-training a natural language processing model using the historical document data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating document drafts using the fine-tuning natural language processing model, means for providing the generated document drafts to users, means for tokenizing the content of the user's request, means for inputting the tokenized data into the natural language processing model, and means for decoding the tokenized data generated from the natural language processing model. This makes it possible for users to easily and efficiently create high-quality documents.

[0441] "Past document data" refers to information about previously created documents extracted from a company's internal document database.

[0442] A "natural language processing model" is a machine learning model that analyzes and understands text data, and in particular, it uses deep learning techniques to learn the meaning of sentences.

[0443] "Pre-training" is the process of training a natural language processing model using a large amount of text data to learn basic language patterns and structures.

[0444] "Fine-tuning learning" is the process of setting specific parameters for a pre-trained natural language processing model and making adjustments to suit a more specific context.

[0445] "Specific parameters" refer to values ​​or conditions (e.g., amount, time) that are important elements when generating a document.

[0446] A "user request" refers to instructions or requests regarding document generation that a user sends to the system via their device.

[0447] "Tokenization" is the process of dividing text data into smaller units (tokens) such as words and phrases.

[0448] "Decoding" is the process of converting tokenized data back into natural language text in a format that is easy for humans to read.

[0449] A "draft document" is a preliminary version of a document generated by a natural language processing model and created based on user requests.

[0450] This invention is a system for efficiently creating documents within a company. It collects past document data, pre-trains a natural language processing model using that data, and then performs fine-tuning learning using specific parameters, thereby automatically generating document drafts based on user requests.

[0451] First, the server collects historical document data from the company's internal database. This data includes various types of documents such as approval forms, reports, and contracts. The collected data is saved in text or CSV format in preparation for the next processing step.

[0452] Next, the server prepares a natural language processing model (e.g., BERT) using the Python Torch library. It loads historical document data and uses this data to pre-train the model. Specifically, it tokenizes the document data using a tokenizer (e.g., BertTokenizer) and inputs this tokenized data into the model. This allows the model to learn document patterns and vocabulary.

[0453] Once pre-training is complete, the server fine-tunes the model using specific parameters (e.g., amount or time period). This fine-tuning is also performed by tokenizing data using a tokenizer and feeding it to the model. This allows the model to generate more specific and accurate document drafts adapted to the particular business context.

[0454] The user sends a request to the chatbot to create an approval document using their device. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023"). The device then sends this request to the server.

[0455] The server tokenizes the received request and feeds it to a finely tuned model. The model generates a draft document based on the request. The generated tokenized data is decoded again into a readable draft document. Finally, the terminal displays this draft document to the user.

[0456] Specific example

[0457] For example, if a user sends a request to the chatbot to create an approval document for a payment of 5 million yen on December 15, 2023, the following process will occur:

[0458] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0459] 2. The device sends this request to the server.

[0460] 3. The server tokenizes the request and feeds it to a finely tuned natural language processing model.

[0461] 4. The server generates a draft approval document using a natural language processing model.

[0462] 5. The server decodes the generated content and creates a draft document stating, "Proposed content: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023."

[0463] 6. The terminal displays this draft document to the user.

[0464] This system allows users to obtain appropriate and consistent document drafts in a short amount of time, thereby improving work efficiency.

[0465] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0466] Step 1:

[0467] The server executes a mechanism to collect historical document data from the company's internal database. It uses database queries to retrieve document data in text or CSV format, such as approval forms, reports, and contracts. The database queries used as input output document data, which is saved to the "past_documents.csv" file. Specifically, it calls an API for database access and executes SQL queries.

[0468] Step 2:

[0469] The server prepares and configures a natural language processing model (e.g., BERT). It loads the model using the Python Torch library and sets the tokenizer, for example, `tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')`. It reads a past document data file, "past_documents.csv", as input and tokenizes it using the tokenizer. The output is the tokenized document data.

[0470] Step 3:

[0471] The server pre-trains a BERT model using past document data. The tokenized data from the tokenizer is created as a dataset using `dataset = TextDataset(tokenizer, 'past_documents.csv')`, and the model is pre-trained using `model.train(dataset)`. The pre-trained model is then output using the tokenized data output from the tokenizer as input.

[0472] Step 4:

[0473] The server performs fine-tuning on a pre-trained model using specific parameters (e.g., amount or time period). It sets specific parameters, tokenizes the fine-tuning data using a tokenizer, and feeds it to the model. Specifically, it executes `model.fine_tune(tokenized_data, parameters)`. The input is the specific parameters and tokenized fine-tuning data, and the output is the fine-tuned model.

[0474] Step 5:

[0475] The user sends a request for the creation of an approval document to the chatbot via their device. Specifically, they enter "Amount: 5 million yen, Date: December 15, 2023" into the chatbot's interface and press the send button. The input is the user's prompt text, and the output is the request sent by the chatbot's interface.

[0476] Step 6:

[0477] The terminal sends the user's request to the server. Specifically, it creates an HTTP POST request and sends the request data to a specific endpoint on the server (e.g., " / generate_document"). The input is the user's request, and the output is the sending of the HTTP request to the server.

[0478] Step 7:

[0479] The server tokenizes the request content and provides it to a fine-tuned model. The received request is tokenized using `tokenizer.tokenize(request_data)`, and a draft document is generated using `model.generate(tokenized_request)`. The input is the tokenized request data, and the output is the generated draft document.

[0480] Step 8:

[0481] The server decodes the generated tokenized data and creates a draft document. The token output is decoded using tokenizer.decode(generated_tokens) and converted into a user-readable format. The input is the tokenized draft document data, and the output is the decoded draft document.

[0482] Step 9:

[0483] The terminal displays the generated document draft to the user. Specifically, it receives the document draft from the server and displays it on the chatbot screen. The input is the document draft from the server, and the output is the display to the user.

[0484] (Application Example 1)

[0485] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0486] Logistics centers require a wide variety of documents (purchase orders, shipping instructions, inventory management reports, etc.), but creating these documents is time-consuming and labor-intensive. Furthermore, because these documents must accurately reflect past data and current conditions, they are prone to human error. While efficiency is needed in these document creation processes, on-site personnel often require advanced specialized knowledge. A system is needed to address these challenges.

[0487] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0488] In this invention, the server includes means for collecting historical data, means for pre-training a natural language processing model using the historical data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating draft documents using the fine-tuning natural language processing model, means for providing the generated draft documents to users, means for collecting historical order history and inventory data from the logistics center's database, means for automatically generating purchase orders, shipping instructions, and inventory management reports using the natural language processing model, and means for displaying and modifying the generated draft documents on a smartphone. This makes it possible to automatically generate documents efficiently and accurately at the logistics center, significantly reducing the burden on on-site staff and improving work efficiency.

[0489] "Past data" refers to information including past order history and inventory data at the logistics center.

[0490] A "natural language processing model" is a machine learning model that understands human language and generates documents.

[0491] "Pre-training" is the process of training a natural language processing model from the ground up using a large amount of data.

[0492] "Fine-tuning learning" is the process of optimizing a natural language processing model, which has been pre-trained using specific parameters, to suit a particular application.

[0493] A "user request" is information that a user uses to request document generation based on specific parameters.

[0494] A "draft document" is a document generated by a natural language processing model and submitted as a candidate.

[0495] "Means of provision" refers to the method by which the generated draft document is displayed for the user to review and modify.

[0496] A "logistics center database" is a data system that stores order history and inventory data managed within a logistics center.

[0497] A "purchase order" is a document that details the order for goods or services.

[0498] A "shipping instruction sheet" is a document that contains instructions for shipping goods out of a logistics center.

[0499] An "inventory management report" is a report that summarizes the current inventory status and related information.

[0500] A "smartphone" is a portable, multi-functional device that can display and process information using applications.

[0501] This invention is a system for efficiently and automatically generating documents in a logistics center. This system collects historical data, uses that data to pre-train a natural language processing model, and then fine-tunes it using specific parameters. Next, it generates draft documents based on user requests and provides them to the user.

[0502] The specific configuration of this system is as follows:

[0503] Hardware and software:

[0504] Server: Responsible for data collection, model training, fine-tuning, and document generation. The natural language processing model used includes BERT.

[0505] Database system: Stores order history and inventory data for the logistics center. Uses MySQL, PostgreSQL, etc.

[0506] Tokenizer: Uses a Python library (such as Hugging Face's Transformers) to tokenize the collected data.

[0507] User interface device (smartphone): Built with React Native to display generated document drafts and receive user requests.

[0508] Program processing:

[0509] 1. Data collection:

[0510] The server collects past order history and inventory data from the logistics center's database. The data is saved in text file or CSV format.

[0511] 2. Pre-learning:

[0512] The server sets up a natural language processing model (BERT) and pre-trains the model using collected historical data. The data is tokenized using the Python Hugging Face library and input into the model.

[0513] 3. Fine-tuning learning:

[0514] The server fine-tunes a pre-trained model using specific parameters (product name, quantity, delivery date). This process also uses a tokenizer to tokenize the data and add optimization information to the model.

[0515] 4. Receiving and sending user requests:

[0516] Users use their smartphones to submit requests for new purchase orders or shipping orders. These requests include the necessary parameters (e.g., "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023").

[0517] The smartphone sends this request to the server.

[0518] 5. Document draft generation:

[0519] The server tokenizes the received request and feeds it into a finely tuned model. The model generates a draft document based on the specified parameters.

[0520] 6. Providing draft documents:

[0521] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. For example, it might generate a specific draft document such as, "Please deliver 100 units of product A on December 1, 2023."

[0522] The generated document draft is displayed to the user on their smartphone. The user can review the content and make corrections as needed.

[0523] This will enable the automation and streamlining of document generation tasks in logistics centers. Furthermore, it will allow users to obtain accurate documents in a shorter time, improving both work efficiency and accuracy.

[0524] Specific example:

[0525] When a user submits a request to create a purchase order with the following information: "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023," the server tokenizes this information and feeds it into a finely tuned natural language processing model to generate a purchase order like this: "Please deliver 100 units of Product A on December 1, 2023." This draft document is then displayed to the user on their smartphone.

[0526] Example of a prompt:

[0527] "Please create a purchase order using the following parameters: Product Name: Product A, Quantity: 100, Due Date: December 1, 2023"

[0528] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0529] Step 1:

[0530] The server collects past order history and inventory data from the logistics center's database. It queries the database system (e.g., MySQL or PostgreSQL) to extract the necessary data. The collected data is saved in text file or CSV format.

[0531] Input: Logistics center database

[0532] Output: Text or CSV file containing past order history and inventory data.

[0533] Step 2:

[0534] The server pre-trains a natural language processing model (BERT) using the collected historical data. The data is tokenized using the Python Hugging Face library and input into the model. This allows the model to learn basic language structures and patterns.

[0535] Input: Text or CSV file containing past order history and inventory data.

[0536] Output: Pre-trained natural language processing model

[0537] Step 3:

[0538] The server fine-tunes a pre-trained model using specific parameters (product name, quantity, delivery date). The Python Hugging Face library is used to tokenize additional data and add optimization information to the model. This allows the model to adapt to specific business contexts.

[0539] Input: Specific parameters (product name, quantity, delivery date)

[0540] Output: Fine-tuned natural language processing model

[0541] Step 4:

[0542] Users use their smartphones to submit requests for new purchase orders or shipping orders. These requests include the necessary parameters (e.g., "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023").

[0543] Input: Request parameters (product name, quantity, delivery date)

[0544] Output: Request from the user's smartphone

[0545] Step 5:

[0546] The terminal sends the user's request to the server. The request data is transferred to the server using protocols such as HTTP requests.

[0547] Input: Request from the user's smartphone

[0548] Output: Request data to the server

[0549] Step 6:

[0550] The server tokenizes the received request using a tokenizer and feeds it into a finely tuned model. The model generates a draft document based on the specified parameters.

[0551] Input: Request data to the server

[0552] Output: Tokenized data of the generated document

[0553] Step 7:

[0554] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Specifically, it converts the tokenized data into natural language sentences and formats it as a structured document.

[0555] Input: Tokenized data of the generated document

[0556] Output: Draft document in an easy-to-read format

[0557] Step 8:

[0558] The terminal displays the generated document draft to the user. The user can review the document draft through the smartphone interface and make revisions as needed.

[0559] Input: Draft document in an easy-to-read format

[0560] Output: Draft document that users can review and edit.

[0561] The above steps automate and streamline document generation at the logistics center. Clearly defining the specific actions and data input / output performed at each step makes system implementation and operation easier.

[0562] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0563] The system of the present invention automatically generates draft documents based on user requests by collecting past document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the content and tone of the draft documents can be adjusted according to the user's emotions. This system is implemented according to the following procedure.

[0564] Program processing

[0565] 1. The server collects historical document data from the company's internal database. The collected data is saved in text file or CSV format.

[0566] 2. The server sets up a natural language processing model (e.g., BERT). The collected document data is tokenized using a tokenizer and input into the model. The model is then pre-trained to learn common document patterns and vocabulary.

[0567] 3. The server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning allows the model to adapt to the specific business context and improves its accuracy.

[0568] 4. The user sends a request to the chatbot from their device to create a proposal document. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0569] 5. The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes emotions from the input text or voice and sends the results to the server.

[0570] 6. The server tokenizes the user's request based on the sentiment analysis results received from the sentiment engine and inputs it into a finely tuned model. The model generates a draft document taking the sentiment analysis results into account.

[0571] 7. The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Based on the sentiment analysis results, the tone and content of the document are adjusted.

[0572] 8. The terminal displays the generated draft document to the user. This allows the user to quickly obtain a draft document that takes into account an appropriate and emotionally balanced tone.

[0573] Specific example

[0574] For example, if a user submits a request to create a proposal document for a transaction with the following conditions: "Amount: 5 million yen, Date: December 15, 2023, Emotion: Angry," the following process will occur:

[0575] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0576] 2. The device sends this request to the emotion engine to analyze the user's emotions. If the analysis result is "anger," the tone is taken into consideration.

[0577] 3. The device sends the request to the server along with the sentiment analysis results.

[0578] 4. The server tokenizes the request and feeds it into a finely tuned natural language processing model. The model generates a draft document, taking sentiment analysis results into account.

[0579] 5. The server decodes the draft document and generates content such as, "Proposal: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[0580] 6. The terminal displays this draft document to the user.

[0581] This series of processes allows users to efficiently create approval documents that take emotions into consideration, thereby improving the accuracy and effectiveness of communication.

[0582] The following describes the processing flow.

[0583] Step 1:

[0584] The server collects historical document data from the company's internal database. The collected data is saved in text file or CSV format.

[0585] Step 2:

[0586] The server configures a natural language processing model (e.g., BERT). It tokenizes the collected document data using a tokenizer. The tokenizer divides the document data into units of words and phrases, converting it into a format that the model can understand.

[0587] Step 3:

[0588] The server inputs tokenized data into a natural language processing model for pre-training. The model learns common document patterns and vocabulary from past document data.

[0589] Step 4:

[0590] The server performs fine-tuning learning on a pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning allows the model to adapt to a specific business context, improving its accuracy and adaptability.

[0591] Step 5:

[0592] The user sends a request to the chatbot from their device to create an approval document. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0593] Step 6:

[0594] The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the emotions from the input text or voice and sends the results to the server.

[0595] Step 7:

[0596] The server tokenizes the user's request based on the sentiment analysis results received from the sentiment engine and inputs it into a finely tuned model. The model then generates a draft document, taking the sentiment analysis results into account.

[0597] Step 8:

[0598] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. The generated draft document is then adjusted in tone and content based on the sentiment analysis results.

[0599] Step 9:

[0600] The terminal displays the generated draft document to the user. This allows the user to quickly review a draft document that takes into account appropriate and emotional tone.

[0601] Specific example

[0602] For example, if a user submits a request to create a proposal document for a transaction with the following conditions: "Amount: 5 million yen, Date: December 15, 2023, Emotion: Angry," the following process will occur:

[0603] Step 1:

[0604] The user submits a request from their device to create an approval document for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0605] Step 2:

[0606] The device sends this request to the emotion engine, which analyzes the user's emotions. The analysis results detect that the user is "angry."

[0607] Step 3:

[0608] The emotion engine sends the emotion analysis results to the server.

[0609] Step 4:

[0610] The server tokenizes the request along with the sentiment analysis results and inputs them into a finely tuned model. The model then generates a draft document, taking the sentiment analysis results into account.

[0611] Step 5:

[0612] The server decodes the generated draft document and produces content such as, "Proposal: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[0613] Step 6:

[0614] The terminal displays the generated draft document to the user.

[0615] In this way, users can create emotionally conscious document drafts quickly and efficiently.

[0616] (Example 2)

[0617] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0618] Conventional document generation systems failed to consider user emotions, resulting in a uniform tone in the generated documents. This led to the problem of documents not adequately reflecting user intent and feelings. Furthermore, methods for providing highly accurate models adapted to business contexts through fine-tuning learning using past document data were insufficient.

[0619] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0620] In this invention, the server includes means for collecting past document data, means for pre-training a natural language processing model using the past document data, means for fine-tuning the natural language processing model using specific parameters, means for analyzing user requests to recognize emotional states, and means for adjusting the tone and content of draft documents considering the emotional states. This makes it possible to generate documents with an appropriate tone that reflects the user's emotions. Furthermore, it is possible to provide highly accurate documents adapted to a specific business context.

[0621] "Past document data" refers to documents previously created by companies or individuals that are stored in databases or file systems.

[0622] A "natural language processing model" is an artificial intelligence model used to understand and generate human language, and includes models such as BERT.

[0623] "Pre-training" is the process of initially training a model using a large amount of collected data to teach it common language patterns and vocabulary.

[0624] "Fine-tuning learning" is the process of further training a pre-trained model according to specific tasks or contexts to improve the model's accuracy.

[0625] A "user request" is input from a user to give specific instructions or requests to the system, and is sent through a chatbot or form.

[0626] A "document draft" is a draft of a document, such as a proposal or report, generated through user requests and system processing.

[0627] "Emotional state" refers to the state of emotions analyzed from the user's input, and includes emotions such as anger, joy, and sadness.

[0628] "Tokenization" is the process of breaking down text data into units of words or subwords and assigning an ID to each of them.

[0629] An "emotion engine" is software or an algorithm that analyzes a user's text or voice input to recognize their emotional state.

[0630] "Decoding" is the process of converting tokenized data back into a human-readable text format.

[0631] The system of this invention automatically generates draft documents based on user requests by collecting past document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the content and tone of the draft documents can be adjusted according to the user's emotions.

[0632] Specifically, the server first collects historical document data from the company's internal database. This data includes business documents such as sales reports, meeting minutes, and approval documents. The collected data is saved in text file or CSV format.

[0633] Next, the server loads a natural language processing model (e.g., BERT) using the Hugging Face library. The collected document data is tokenized using a tokenizer (e.g., BERT Tokenizer) and input into the model. The model is pre-trained using this tokenized data to learn common language patterns and vocabulary.

[0634] Subsequently, the server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount and timing). This fine-tuning allows the model to adapt to the specific business context, enabling more accurate document generation. The fine-tuned model is stored on the server and used to respond to user requests.

[0635] The user uses a chatbot interface from their device to request the creation of an approval document for specific details (e.g., "Amount: 5 million yen, Date: December 15, 2023"). This request is sent from the device to the emotion engine, which analyzes the user's emotional state based on their input. The emotion engine classifies the user's emotions into categories such as "anger," "joy," and "sadness," and returns the result to the device.

[0636] The terminal sends the sentiment analysis results and the user's request to the server, which tokenizes the request and inputs it into a finely tuned model. The model takes the sentiment analysis results into account and generates a draft document that reflects the user's emotions. For example, if anger is detected, the tone of the document is strengthened and adjusted to convey a request for immediate action.

[0637] The generated tokenized data is decoded on the server and sent to the terminal as a readable, formatted draft document. The terminal displays the final draft document to the user, who then reviews the generated document and makes any necessary corrections or additions.

[0638] As a concrete example, consider a case where a user submits a request to create a proposal document with the following conditions: "Payment amount: 5 million yen, Date: December 15, 2023, Emotion: Angry." In response to this request, the server generates a draft document reflecting the user's emotion and provides it to the user. The generated draft document would be displayed with content such as: "Proposal: The payment amount for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[0639] This system allows users to efficiently create documents that take emotions into account, thereby improving the accuracy and effectiveness of communication.

[0640] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0641] Step 1:

[0642] The server collects historical document data from the company's internal database. This includes business documents such as sales reports, meeting minutes, and approval documents. The collected data is saved in text file or CSV format.

[0643] Input: Corporate database

[0644] Output: Collected historical document data (text files, CSV)

[0645] Step 2:

[0646] The server loads a natural language processing model (e.g., BERT) using the Hugging Face library. It tokenizes the collected document data using a tokenizer (e.g., BERT Tokenizer). Each document is broken down into words and subwords, and converted into integer IDs.

[0647] Input: Past document data

[0648] Output: Tokenized document data

[0649] Step 3:

[0650] The server inputs tokenized data into the model and performs pre-training. During this process, the model learns document patterns and vocabulary. An initially trained model is obtained.

[0651] Input: Tokenized document data

[0652] Output: Pre-trained natural language processing model

[0653] Step 4:

[0654] The server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount or timing). This allows the model to adapt to the specific business context and improve its accuracy.

[0655] Input: Pre-trained natural language processing model, specific parameters

[0656] Output: Fine-tuned natural language processing model

[0657] Step 5:

[0658] The user sends a request for the creation of an approval document using the chatbot interface from their device. The request includes specific parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0659] Input: User request (payment amount, timing, etc.)

[0660] Output: Request data received by the terminal

[0661] Step 6:

[0662] The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine classifies the emotional state from the input text or voice and returns the result to the device.

[0663] Input: User request, emotion engine

[0664] Output: Emotion analysis results (anger, joy, sadness, etc.)

[0665] Step 7:

[0666] The device sends the sentiment analysis results and the user's request to the server.

[0667] Input: Sentiment analysis results, user requests

[0668] Output: Sentiment analysis results and requests received by the server

[0669] Step 8:

[0670] The server tokenizes the request and inputs it into a finely tuned model. The model takes sentiment analysis results into account and generates a draft document that reflects the user's emotions.

[0671] Input: Tokenized request content, sentiment analysis results, and a fine-tuned natural language processing model.

[0672] Output: Generated draft document

[0673] Step 9:

[0674] The server decodes the generated tokenized data and creates a readable, formatted draft document. Based on the sentiment analysis results, the tone and content of the document are adjusted.

[0675] Input: Generated tokenized data

[0676] Output: Formatted draft document

[0677] Step 10:

[0678] The terminal displays the final draft document to the user. The user can review the generated document and make corrections or additions as needed.

[0679] Input: Formatted draft document

[0680] Output: Draft document displayed to the user

[0681] (Application Example 2)

[0682] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0683] Traditional customer service systems have struggled to analyze customer emotions in real time and respond accordingly. Furthermore, relying on manuals makes it difficult to improve customer satisfaction. A system is needed to solve these problems and provide more advanced customer service.

[0684] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0685] In this invention, the server includes means for collecting past document data, means for pre-training a natural language processing model using the past document data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating draft documents using the fine-tuning natural language processing model, means for providing the generated draft documents to the user, and means for analyzing the customer's emotions from their statements and facial expressions, generating a draft document based on the analysis results, and displaying it on a display device. This makes it possible to analyze the customer's emotions in real time and generate and provide an appropriate draft document based on the results.

[0686] "Past document data" refers to existing document information collected from within or outside the company, which is used for pre-training and analysis of the model.

[0687] A "natural language processing model" is a machine learning model designed to enable human understanding of natural language, and is a technology for automatically generating, analyzing, translating, and summarizing document content.

[0688] "Pre-training" is the process of using a broad dataset to teach a model basic language patterns and vocabulary before fine-tuning it for a specific task.

[0689] "Fine-tuning" is an additional learning process that uses specific parameters to further adapt a pre-trained model to a particular task or context.

[0690] A "user request" refers to the specific requests and parameters that a user provides to the system, and is the input information that the system uses to generate appropriate output based on these requests.

[0691] "Customer statements" refer to the words and sentences that customers make to the system, and by analyzing them, we can understand their intentions and requirements.

[0692] "Customer facial expressions" refer to the emotions and reactions shown by a customer's facial expressions, and by analyzing these, information is used to infer the customer's current emotional state.

[0693] "Emotional analysis" refers to the process of automatically identifying customer emotions from text and image data, and then generating appropriate responses and actions based on those results.

[0694] "Generating document drafts" means automatically creating appropriate document content using a natural language processing model based on user requests and analysis results.

[0695] "Displaying on a display device" means displaying information on a device that provides users with a visual representation of the generated document draft or analysis results, such as smart glasses or a display screen.

[0696] The present invention provides a system for pre-training a natural language processing model based on past document data, generating draft documents based on user requests and sentiment analysis results, and displaying them on a display device. The configuration and method for implementing this system are described below.

[0697] System Configuration

[0698] The system of the present invention consists of the following main elements:

[0699] 1. Server

[0700] Data collection method: Collect and store historical document data.

[0701] Natural language processing model: Pre-training is performed using collected document data, and then fine-tuning is performed using specific parameters (e.g., amount, timing).

[0702] Emotion analysis engine: Analyzes customer emotions from their statements and facial expressions.

[0703] 2. Terminal

[0704] Input method: Receives requests from users.

[0705] Display method: The generated document draft is displayed using smart glasses or a display.

[0706] Data collection and learning

[0707] The server collects historical document data from the company's database. This data is stored in text file or CSV format. The collected data is used to pre-train a natural language processing model (e.g., BERT). During pre-training, a tokenizer is used to tokenize the data, which is then input into the model. Next, the pre-trained model is fine-tuned with specific parameters.

[0708] Sentiment analysis and document generation

[0709] The user sends a request through a device (e.g., smart glasses). The request includes parameters such as "Payment amount: 5 million yen, Date: December 15, 2023". The device sends the user's speech and facial expressions to an emotion analysis engine to analyze the customer's emotions. Once the emotion analysis results are obtained, they are sent to the server.

[0710] The server generates a draft document using a finely tuned natural language processing model based on the received sentiment analysis results and request content. During this process, it adjusts the tone and content of the document, taking the sentiment analysis results into consideration. The generated draft document is then sent back to the terminal and displayed on the display device.

[0711] Hardware and software to use

[0712] 1. Hardware:

[0713] Smart glasses: Display customer information and analysis results in real time.

[0714] Webcam: Used to analyze customer facial expressions.

[0715] 2. Software:

[0716] Natural language processing model: BERT is used to analyze user utterances.

[0717] Sentiment analysis engine: Uses the TextBlob library to analyze emotions from customer statements.

[0718] OpenCV: Used to analyze customer facial expressions from image data.

[0719] Specific example

[0720] For example, if a customer says, "I'm having trouble with this product," the following process will occur:

[0721] 1. The user (store clerk) wears smart glasses and interacts with the customer.

[0722] 2. The user's spoken content and facial expression data captured by the webcam are sent to the emotion analysis engine, which then analyzes their emotions.

[0723] 3. The sentiment analysis results and request details are sent to the server, and a draft document is generated.

[0724] 4. The generated draft document is displayed on the smart glasses' screen.

[0725] Examples of prompts to input into a generative AI model include:

[0726] A customer asked a question about our company's products. They seemed confused, so please generate a reassuring message by providing FAQs and relevant product information.

[0727] This method allows users to obtain draft documents in real time that take into account appropriate and emotional tone, enabling them to provide customers with a more satisfying service.

[0728] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0729] Step 1:

[0730] The server collects historical document data from the company's database. This includes data in text file and CSV format. The collected document data is used as input to pre-train a natural language processing model.

[0731] Step 2:

[0732] The server tokenizes the collected document data using a tokenizer and inputs it into a natural language processing model (e.g., BERT). The model is then pre-trained to learn common document patterns and vocabulary. The output is the pre-trained model.

[0733] Step 3:

[0734] The server performs fine-tuning training on a pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning adapts the model to a specific business context, improving its accuracy. The output is the fine-tuned model.

[0735] Step 4:

[0736] The user sends a request from their device. The request includes the necessary parameters (e.g., "Payment amount: 5 million yen, Date: December 15, 2023"). The device sends the user's request to the sentiment analysis engine. The input is the user's request, and the output is the data transfer to the sentiment analysis engine.

[0737] Step 5:

[0738] The device analyzes the user's speech and facial expressions using an emotion analysis engine. The emotion analysis engine analyzes emotions from input text and image data and sends the results to the server. The input is the user's speech and facial expression data, and the output is the emotion analysis result.

[0739] Step 6:

[0740] The server inputs the generated document draft, based on the sentiment analysis results and the user's request, into a finely tuned natural language processing model. The model generates the document draft taking the sentiment analysis results into account. The input is the sentiment analysis results and the user's request, and the output is the generated document draft.

[0741] Step 7:

[0742] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Based on the sentiment analysis results, the tone and content of the document are adjusted. The input is the generated tokenized data, and the output is the draft document in an easy-to-read format.

[0743] Step 8:

[0744] The terminal displays the generated document draft on a display device (e.g., smart glasses or a screen). This allows the user to quickly obtain a document draft that is appropriate and considers emotional tone. The input is a readable document draft, and the output is a display on a display device.

[0745] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0746] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0747] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0748] [Third Embodiment]

[0749] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0750] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0751] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0752] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0753] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0754] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0755] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0756] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0757] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0758] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0759] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0760] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0761] The system of the present invention automatically generates draft documents based on user requests by collecting historical document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. This system is implemented according to the following procedure.

[0762] Program processing

[0763] 1. The server first collects historical document data from the company's internal database. This includes various official documents, such as approval forms. The collected data is saved in text file or CSV format.

[0764] 2. The server then sets up a natural language processing model (e.g., BERT). It loads the collected historical document data and uses this data to pre-train the model. It tokenizes the document data using a tokenizer and feeds it to the model as input to learn patterns and vocabulary.

[0765] 3. The server performs fine-tuning learning on the pre-trained model using specific parameters (e.g., amount or timing). This allows the model to adapt to a specific business context and generate more specific and accurate document drafts. Fine-tuning learning is also performed by tokenizing data using a tokenizer and feeding it to the model.

[0766] 4. The user sends a request to the chatbot to create an approval document using their own device. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0767] 5. The terminal sends the user's request to the server. The server receives this request, tokenizes the request content, and provides it to the refined model.

[0768] 6. The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Specifically, the content generated by the model is processed into a draft document and converted into a format that the user can review.

[0769] 7. The terminal displays the generated document draft to the user. This allows the user to obtain a suitable document draft in a short amount of time.

[0770] Specific example

[0771] For example, if a user sends a request to the chatbot to create an approval document for a payment of 5 million yen on December 15, 2023, the following process will occur:

[0772] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0773] 2. The device sends this request to the server.

[0774] 3. The server tokenizes the request and feeds it to a finely tuned natural language processing model.

[0775] 4. The server generates a draft approval document and decodes it as follows: "Proposal: The budget for this project is 5 million yen, and the implementation date is December 15, 2023."

[0776] 5. The terminal displays this draft document to the user.

[0777] In this way, the system of the present invention provides an environment in which users can easily create approval documents, thereby reducing time and effort.

[0778] The following describes the processing flow.

[0779] Step 1:

[0780] The server collects historical document data from the company's internal database. This includes data from official documents such as approval requests and reports. The collected data is stored in a single text file or a CSV file.

[0781] Step 2:

[0782] The server initializes a natural language processing model (e.g., BERT). This involves using a tool called a tokenizer to tokenize the collected text data (divide it into units of words or sentences) and convert it into a format that the model can understand.

[0783] Step 3:

[0784] The server inputs tokenized data into a natural language processing model for pre-training. This allows the model to learn common document patterns and vocabulary from past document data.

[0785] Step 4:

[0786] The server prepares specific parameters (e.g., amount and time period). Based on these specific parameters, it further refines the pre-trained model. This improves the model's accuracy and adaptability to the specific business context.

[0787] Step 5:

[0788] The user sends a request to the chatbot from their device to create an approval document. This request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0789] Step 6:

[0790] The terminal sends data to the server via a communication interface in order to send user requests to the server.

[0791] Step 7:

[0792] The server tokenizes the user's request using a tokenizer. This tokenized data is then input into a finely tuned natural language processing model.

[0793] Step 8:

[0794] The server decodes the tokenized data generated from the model and creates a draft document in an easy-to-read format. Specifically, it converts the generated tokenized data into natural language sentences.

[0795] Step 9:

[0796] The terminal displays the draft document sent from the server to the user. This allows the user to quickly review an appropriate draft document.

[0797] This series of processes allows users to efficiently create approval documents, significantly reducing the time and effort required.

[0798] (Example 1)

[0799] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0800] In modern businesses, document creation is a time-consuming and labor-intensive task. In particular, creating new documents based on past documents is cumbersome, requiring careful reference and editing, and maintaining consistency is difficult. Furthermore, generating draft documents quickly and accurately in response to specific user requests is challenging to do manually. This leads to a decrease in overall business efficiency.

[0801] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0802] In this invention, the server includes means for collecting historical document data, means for pre-training a natural language processing model using the historical document data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating document drafts using the fine-tuning natural language processing model, means for providing the generated document drafts to users, means for tokenizing the content of the user's request, means for inputting the tokenized data into the natural language processing model, and means for decoding the tokenized data generated from the natural language processing model. This makes it possible for users to easily and efficiently create high-quality documents.

[0803] "Past document data" refers to information about previously created documents extracted from a company's internal document database.

[0804] A "natural language processing model" is a machine learning model that analyzes and understands text data, and in particular, it uses deep learning techniques to learn the meaning of sentences.

[0805] "Pre-training" is the process of training a natural language processing model using a large amount of text data to learn basic language patterns and structures.

[0806] "Fine-tuning learning" is the process of setting specific parameters for a pre-trained natural language processing model and making adjustments to suit a more specific context.

[0807] "Specific parameters" refer to values ​​or conditions (e.g., amount, time) that are important elements when generating a document.

[0808] A "user request" refers to instructions or requests regarding document generation that a user sends to the system via their device.

[0809] "Tokenization" is the process of dividing text data into smaller units (tokens) such as words and phrases.

[0810] "Decoding" is the process of converting tokenized data back into natural language text in a format that is easy for humans to read.

[0811] A "draft document" is a preliminary version of a document generated by a natural language processing model and created based on user requests.

[0812] This invention is a system for efficiently creating documents within a company. It collects past document data, pre-trains a natural language processing model using that data, and then performs fine-tuning learning using specific parameters, thereby automatically generating document drafts based on user requests.

[0813] First, the server collects historical document data from the company's internal database. This data includes various types of documents such as approval forms, reports, and contracts. The collected data is saved in text or CSV format in preparation for the next processing step.

[0814] Next, the server prepares a natural language processing model (e.g., BERT) using the Python Torch library. It loads historical document data and uses this data to pre-train the model. Specifically, it tokenizes the document data using a tokenizer (e.g., BertTokenizer) and inputs this tokenized data into the model. This allows the model to learn document patterns and vocabulary.

[0815] Once pre-training is complete, the server fine-tunes the model using specific parameters (e.g., amount or time period). This fine-tuning is also performed by tokenizing data using a tokenizer and feeding it to the model. This allows the model to generate more specific and accurate document drafts adapted to the particular business context.

[0816] The user sends a request to the chatbot to create an approval document using their device. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023"). The device then sends this request to the server.

[0817] The server tokenizes the received request and feeds it to a finely tuned model. The model generates a draft document based on the request. The generated tokenized data is decoded again into a readable draft document. Finally, the terminal displays this draft document to the user.

[0818] Specific example

[0819] For example, if a user sends a request to the chatbot to create an approval document for a payment of 5 million yen on December 15, 2023, the following process will occur:

[0820] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0821] 2. The device sends this request to the server.

[0822] 3. The server tokenizes the request and feeds it to a finely tuned natural language processing model.

[0823] 4. The server generates a draft approval document using a natural language processing model.

[0824] 5. The server decodes the generated content and creates a draft document stating, "Proposed content: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023."

[0825] 6. The terminal displays this draft document to the user.

[0826] This system allows users to obtain appropriate and consistent document drafts in a short amount of time, thereby improving work efficiency.

[0827] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0828] Step 1:

[0829] The server executes a mechanism to collect historical document data from the company's internal database. It uses database queries to retrieve document data in text or CSV format, such as approval forms, reports, and contracts. The database queries used as input output document data, which is saved to the "past_documents.csv" file. Specifically, it calls an API for database access and executes SQL queries.

[0830] Step 2:

[0831] The server prepares and configures a natural language processing model (e.g., BERT). It loads the model using the Python Torch library and sets the tokenizer, for example, `tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')`. It reads a past document data file, "past_documents.csv", as input and tokenizes it using the tokenizer. The output is the tokenized document data.

[0832] Step 3:

[0833] The server pre-trains a BERT model using past document data. The tokenized data from the tokenizer is created as a dataset using `dataset = TextDataset(tokenizer, 'past_documents.csv')`, and the model is pre-trained using `model.train(dataset)`. The pre-trained model is then output using the tokenized data output from the tokenizer as input.

[0834] Step 4:

[0835] The server performs fine-tuning on a pre-trained model using specific parameters (e.g., amount or time period). It sets specific parameters, tokenizes the fine-tuning data using a tokenizer, and feeds it to the model. Specifically, it executes `model.fine_tune(tokenized_data, parameters)`. The input is the specific parameters and tokenized fine-tuning data, and the output is the fine-tuned model.

[0836] Step 5:

[0837] The user sends a request for the creation of an approval document to the chatbot via their device. Specifically, they enter "Amount: 5 million yen, Date: December 15, 2023" into the chatbot's interface and press the send button. The input is the user's prompt text, and the output is the request sent by the chatbot's interface.

[0838] Step 6:

[0839] The terminal sends the user's request to the server. Specifically, it creates an HTTP POST request and sends the request data to a specific endpoint on the server (e.g., " / generate_document"). The input is the user's request, and the output is the sending of the HTTP request to the server.

[0840] Step 7:

[0841] The server tokenizes the request content and provides it to a fine-tuned model. The received request is tokenized using `tokenizer.tokenize(request_data)`, and a draft document is generated using `model.generate(tokenized_request)`. The input is the tokenized request data, and the output is the generated draft document.

[0842] Step 8:

[0843] The server decodes the generated tokenized data and creates a draft document. The token output is decoded using tokenizer.decode(generated_tokens) and converted into a user-readable format. The input is the tokenized draft document data, and the output is the decoded draft document.

[0844] Step 9:

[0845] The terminal displays the generated document draft to the user. Specifically, it receives the document draft from the server and displays it on the chatbot screen. The input is the document draft from the server, and the output is the display to the user.

[0846] (Application Example 1)

[0847] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0848] Logistics centers require a wide variety of documents (purchase orders, shipping instructions, inventory management reports, etc.), but creating these documents is time-consuming and labor-intensive. Furthermore, because these documents must accurately reflect past data and current conditions, they are prone to human error. While efficiency is needed in these document creation processes, on-site personnel often require advanced specialized knowledge. A system is needed to address these challenges.

[0849] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0850] In this invention, the server includes means for collecting historical data, means for pre-training a natural language processing model using the historical data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating draft documents using the fine-tuning natural language processing model, means for providing the generated draft documents to users, means for collecting historical order history and inventory data from the logistics center's database, means for automatically generating purchase orders, shipping instructions, and inventory management reports using the natural language processing model, and means for displaying and modifying the generated draft documents on a smartphone. This makes it possible to automatically generate documents efficiently and accurately at the logistics center, significantly reducing the burden on on-site staff and improving work efficiency.

[0851] "Past data" refers to information including past order history and inventory data at the logistics center.

[0852] A "natural language processing model" is a machine learning model that understands human language and generates documents.

[0853] "Pre-training" is the process of training a natural language processing model from the ground up using a large amount of data.

[0854] "Fine-tuning learning" is the process of optimizing a natural language processing model, which has been pre-trained using specific parameters, to suit a particular application.

[0855] A "user request" is information that a user uses to request document generation based on specific parameters.

[0856] A "draft document" is a document generated by a natural language processing model and submitted as a candidate.

[0857] "Means of provision" refers to the method by which the generated draft document is displayed for the user to review and modify.

[0858] A "logistics center database" is a data system that stores order history and inventory data managed within a logistics center.

[0859] A "purchase order" is a document that details the order for goods or services.

[0860] A "shipping instruction sheet" is a document that contains instructions for shipping goods out of a logistics center.

[0861] An "inventory management report" is a report that summarizes the current inventory status and related information.

[0862] A "smartphone" is a portable, multi-functional device that can display and process information using applications.

[0863] This invention is a system for efficiently and automatically generating documents in a logistics center. This system collects historical data, uses that data to pre-train a natural language processing model, and then fine-tunes it using specific parameters. Next, it generates draft documents based on user requests and provides them to the user.

[0864] The specific configuration of this system is as follows:

[0865] Hardware and software:

[0866] Server: Responsible for data collection, model training, fine-tuning, and document generation. The natural language processing model used includes BERT.

[0867] Database system: Stores order history and inventory data for the logistics center. Uses MySQL, PostgreSQL, etc.

[0868] Tokenizer: Uses a Python library (such as Hugging Face's Transformers) to tokenize the collected data.

[0869] User interface device (smartphone): Built with React Native to display generated document drafts and receive user requests.

[0870] Program processing:

[0871] 1. Data collection:

[0872] The server collects past order history and inventory data from the logistics center's database. The data is saved in text file or CSV format.

[0873] 2. Pre-learning:

[0874] The server sets up a natural language processing model (BERT) and pre-trains the model using collected historical data. The data is tokenized using the Python Hugging Face library and input into the model.

[0875] 3. Fine-tuning learning:

[0876] The server fine-tunes a pre-trained model using specific parameters (product name, quantity, delivery date). This process also uses a tokenizer to tokenize the data and add optimization information to the model.

[0877] 4. Receiving and sending user requests:

[0878] Users use their smartphones to submit requests for new purchase orders or shipping orders. These requests include the necessary parameters (e.g., "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023").

[0879] The smartphone sends this request to the server.

[0880] 5. Document draft generation:

[0881] The server tokenizes the received request and feeds it into a finely tuned model. The model generates a draft document based on the specified parameters.

[0882] 6. Providing draft documents:

[0883] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. For example, it might generate a specific draft document such as, "Please deliver 100 units of product A on December 1, 2023."

[0884] The generated document draft is displayed to the user on their smartphone. The user can review the content and make corrections as needed.

[0885] This will enable the automation and streamlining of document generation tasks in logistics centers. Furthermore, it will allow users to obtain accurate documents in a shorter time, improving both work efficiency and accuracy.

[0886] Specific example:

[0887] When a user submits a request to create a purchase order with the following information: "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023," the server tokenizes this information and feeds it into a finely tuned natural language processing model to generate a purchase order like this: "Please deliver 100 units of Product A on December 1, 2023." This draft document is then displayed to the user on their smartphone.

[0888] Example of a prompt:

[0889] "Please create a purchase order using the following parameters: Product Name: Product A, Quantity: 100, Due Date: December 1, 2023"

[0890] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0891] Step 1:

[0892] The server collects past order history and inventory data from the logistics center's database. It queries the database system (e.g., MySQL or PostgreSQL) to extract the necessary data. The collected data is saved in text file or CSV format.

[0893] Input: Logistics center database

[0894] Output: Text or CSV file containing past order history and inventory data.

[0895] Step 2:

[0896] The server pre-trains a natural language processing model (BERT) using the collected historical data. The data is tokenized using the Python Hugging Face library and input into the model. This allows the model to learn basic language structures and patterns.

[0897] Input: Text or CSV file containing past order history and inventory data.

[0898] Output: Pre-trained natural language processing model

[0899] Step 3:

[0900] The server fine-tunes a pre-trained model using specific parameters (product name, quantity, delivery date). The Python Hugging Face library is used to tokenize additional data and add optimization information to the model. This allows the model to adapt to specific business contexts.

[0901] Input: Specific parameters (product name, quantity, delivery date)

[0902] Output: Fine-tuned natural language processing model

[0903] Step 4:

[0904] Users use their smartphones to submit requests for new purchase orders or shipping orders. These requests include the necessary parameters (e.g., "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023").

[0905] Input: Request parameters (product name, quantity, delivery date)

[0906] Output: Request from the user's smartphone

[0907] Step 5:

[0908] The terminal sends the user's request to the server. The request data is transferred to the server using protocols such as HTTP requests.

[0909] Input: Request from the user's smartphone

[0910] Output: Request data to the server

[0911] Step 6:

[0912] The server tokenizes the received request using a tokenizer and feeds it into a finely tuned model. The model generates a draft document based on the specified parameters.

[0913] Input: Request data to the server

[0914] Output: Tokenized data of the generated document

[0915] Step 7:

[0916] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Specifically, it converts the tokenized data into natural language sentences and formats it as a structured document.

[0917] Input: Tokenized data of the generated document

[0918] Output: Draft document in an easy-to-read format

[0919] Step 8:

[0920] The terminal displays the generated document draft to the user. The user can review the document draft through the smartphone interface and make revisions as needed.

[0921] Input: Draft document in an easy-to-read format

[0922] Output: Draft document that users can review and edit.

[0923] The above steps automate and streamline document generation at the logistics center. Clearly defining the specific actions and data input / output performed at each step makes system implementation and operation easier.

[0924] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0925] The system of the present invention automatically generates draft documents based on user requests by collecting past document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the content and tone of the draft documents can be adjusted according to the user's emotions. This system is implemented according to the following procedure.

[0926] Program processing

[0927] 1. The server collects historical document data from the company's internal database. The collected data is saved in text file or CSV format.

[0928] 2. The server sets up a natural language processing model (e.g., BERT). The collected document data is tokenized using a tokenizer and input into the model. The model is then pre-trained to learn common document patterns and vocabulary.

[0929] 3. The server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning allows the model to adapt to the specific business context and improves its accuracy.

[0930] 4. The user sends a request to the chatbot from their device to create a proposal document. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0931] 5. The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes emotions from the input text or voice and sends the results to the server.

[0932] 6. The server tokenizes the user's request based on the sentiment analysis results received from the sentiment engine and inputs it into a finely tuned model. The model generates a draft document taking the sentiment analysis results into account.

[0933] 7. The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Based on the sentiment analysis results, the tone and content of the document are adjusted.

[0934] 8. The terminal displays the generated draft document to the user. This allows the user to quickly obtain a draft document that takes into account an appropriate and emotionally balanced tone.

[0935] Specific example

[0936] For example, if a user submits a request to create a proposal document for a transaction with the following conditions: "Amount: 5 million yen, Date: December 15, 2023, Emotion: Angry," the following process will occur:

[0937] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0938] 2. The device sends this request to the emotion engine to analyze the user's emotions. If the analysis result is "anger," the tone is taken into consideration.

[0939] 3. The device sends the request to the server along with the sentiment analysis results.

[0940] 4. The server tokenizes the request and feeds it into a finely tuned natural language processing model. The model generates a draft document, taking sentiment analysis results into account.

[0941] 5. The server decodes the draft document and generates content such as, "Proposal: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[0942] 6. The terminal displays this draft document to the user.

[0943] This series of processes allows users to efficiently create approval documents that take emotions into consideration, thereby improving the accuracy and effectiveness of communication.

[0944] The following describes the processing flow.

[0945] Step 1:

[0946] The server collects historical document data from the company's internal database. The collected data is saved in text file or CSV format.

[0947] Step 2:

[0948] The server configures a natural language processing model (e.g., BERT). It tokenizes the collected document data using a tokenizer. The tokenizer divides the document data into units of words and phrases, converting it into a format that the model can understand.

[0949] Step 3:

[0950] The server inputs tokenized data into a natural language processing model for pre-training. The model learns common document patterns and vocabulary from past document data.

[0951] Step 4:

[0952] The server performs fine-tuning learning on a pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning allows the model to adapt to a specific business context, improving its accuracy and adaptability.

[0953] Step 5:

[0954] The user sends a request to the chatbot from their device to create an approval document. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[0955] Step 6:

[0956] The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the emotions from the input text or voice and sends the results to the server.

[0957] Step 7:

[0958] The server tokenizes the user's request based on the sentiment analysis results received from the sentiment engine and inputs it into a finely tuned model. The model then generates a draft document, taking the sentiment analysis results into account.

[0959] Step 8:

[0960] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. The generated draft document is then adjusted in tone and content based on the sentiment analysis results.

[0961] Step 9:

[0962] The terminal displays the generated draft document to the user. This allows the user to quickly review a draft document that takes into account appropriate and emotional tone.

[0963] Specific example

[0964] For example, if a user submits a request to create a proposal document for a transaction with the following conditions: "Amount: 5 million yen, Date: December 15, 2023, Emotion: Angry," the following process will occur:

[0965] Step 1:

[0966] The user submits a request from their device to create an approval document for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[0967] Step 2:

[0968] The device sends this request to the emotion engine, which analyzes the user's emotions. The analysis results detect that the user is "angry."

[0969] Step 3:

[0970] The emotion engine sends the emotion analysis results to the server.

[0971] Step 4:

[0972] The server tokenizes the request along with the sentiment analysis results and inputs them into a finely tuned model. The model then generates a draft document, taking the sentiment analysis results into account.

[0973] Step 5:

[0974] The server decodes the generated draft document and produces content such as, "Proposal: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[0975] Step 6:

[0976] The terminal displays the generated draft document to the user.

[0977] In this way, users can create emotionally conscious document drafts quickly and efficiently.

[0978] (Example 2)

[0979] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0980] Conventional document generation systems failed to consider user emotions, resulting in a uniform tone in the generated documents. This led to the problem of documents not adequately reflecting user intent and feelings. Furthermore, methods for providing highly accurate models adapted to business contexts through fine-tuning learning using past document data were insufficient.

[0981] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0982] In this invention, the server includes means for collecting past document data, means for pre-training a natural language processing model using the past document data, means for fine-tuning the natural language processing model using specific parameters, means for analyzing user requests to recognize emotional states, and means for adjusting the tone and content of draft documents considering the emotional states. This makes it possible to generate documents with an appropriate tone that reflects the user's emotions. Furthermore, it is possible to provide highly accurate documents adapted to a specific business context.

[0983] "Past document data" refers to documents previously created by companies or individuals that are stored in databases or file systems.

[0984] A "natural language processing model" is an artificial intelligence model used to understand and generate human language, and includes models such as BERT.

[0985] "Pre-training" is the process of initially training a model using a large amount of collected data to teach it common language patterns and vocabulary.

[0986] "Fine-tuning learning" is the process of further training a pre-trained model according to specific tasks or contexts to improve the model's accuracy.

[0987] A "user request" is input from a user to give specific instructions or requests to the system, and is sent through a chatbot or form.

[0988] A "document draft" is a draft of a document, such as a proposal or report, generated through user requests and system processing.

[0989] "Emotional state" refers to the state of emotions analyzed from the user's input, and includes emotions such as anger, joy, and sadness.

[0990] "Tokenization" is the process of breaking down text data into units of words or subwords and assigning an ID to each of them.

[0991] An "emotion engine" is software or an algorithm that analyzes a user's text or voice input to recognize their emotional state.

[0992] "Decoding" is the process of converting tokenized data back into a human-readable text format.

[0993] The system of this invention automatically generates draft documents based on user requests by collecting past document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the content and tone of the draft documents can be adjusted according to the user's emotions.

[0994] Specifically, the server first collects historical document data from the company's internal database. This data includes business documents such as sales reports, meeting minutes, and approval documents. The collected data is saved in text file or CSV format.

[0995] Next, the server loads a natural language processing model (e.g., BERT) using the Hugging Face library. The collected document data is tokenized using a tokenizer (e.g., BERT Tokenizer) and input into the model. The model is pre-trained using this tokenized data to learn common language patterns and vocabulary.

[0996] Subsequently, the server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount and timing). This fine-tuning allows the model to adapt to the specific business context, enabling more accurate document generation. The fine-tuned model is stored on the server and used to respond to user requests.

[0997] The user uses a chatbot interface from their device to request the creation of an approval document for specific details (e.g., "Amount: 5 million yen, Date: December 15, 2023"). This request is sent from the device to the emotion engine, which analyzes the user's emotional state based on their input. The emotion engine classifies the user's emotions into categories such as "anger," "joy," and "sadness," and returns the result to the device.

[0998] The terminal sends the sentiment analysis results and the user's request to the server, which tokenizes the request and inputs it into a finely tuned model. The model takes the sentiment analysis results into account and generates a draft document that reflects the user's emotions. For example, if anger is detected, the tone of the document is strengthened and adjusted to convey a request for immediate action.

[0999] The generated tokenized data is decoded on the server and sent to the terminal as a readable, formatted draft document. The terminal displays the final draft document to the user, who then reviews the generated document and makes any necessary corrections or additions.

[1000] As a concrete example, consider a case where a user submits a request to create a proposal document with the following conditions: "Payment amount: 5 million yen, Date: December 15, 2023, Emotion: Angry." In response to this request, the server generates a draft document reflecting the user's emotion and provides it to the user. The generated draft document would be displayed with content such as: "Proposal: The payment amount for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[1001] This system allows users to efficiently create documents that take emotions into account, thereby improving the accuracy and effectiveness of communication.

[1002] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1003] Step 1:

[1004] The server collects historical document data from the company's internal database. This includes business documents such as sales reports, meeting minutes, and approval documents. The collected data is saved in text file or CSV format.

[1005] Input: Corporate database

[1006] Output: Collected historical document data (text files, CSV)

[1007] Step 2:

[1008] The server loads a natural language processing model (e.g., BERT) using the Hugging Face library. It tokenizes the collected document data using a tokenizer (e.g., BERT Tokenizer). Each document is broken down into words and subwords, and converted into integer IDs.

[1009] Input: Past document data

[1010] Output: Tokenized document data

[1011] Step 3:

[1012] The server inputs tokenized data into the model and performs pre-training. During this process, the model learns document patterns and vocabulary. An initially trained model is obtained.

[1013] Input: Tokenized document data

[1014] Output: Pre-trained natural language processing model

[1015] Step 4:

[1016] The server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount or timing). This allows the model to adapt to the specific business context and improve its accuracy.

[1017] Input: Pre-trained natural language processing model, specific parameters

[1018] Output: Fine-tuned natural language processing model

[1019] Step 5:

[1020] The user sends a request for the creation of an approval document using the chatbot interface from their device. The request includes specific parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[1021] Input: User request (payment amount, timing, etc.)

[1022] Output: Request data received by the terminal

[1023] Step 6:

[1024] The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine classifies the emotional state from the input text or voice and returns the result to the device.

[1025] Input: User request, emotion engine

[1026] Output: Emotion analysis results (anger, joy, sadness, etc.)

[1027] Step 7:

[1028] The device sends the sentiment analysis results and the user's request to the server.

[1029] Input: Sentiment analysis results, user requests

[1030] Output: Sentiment analysis results and requests received by the server

[1031] Step 8:

[1032] The server tokenizes the request and inputs it into a finely tuned model. The model takes sentiment analysis results into account and generates a draft document that reflects the user's emotions.

[1033] Input: Tokenized request content, sentiment analysis results, and a fine-tuned natural language processing model.

[1034] Output: Generated draft document

[1035] Step 9:

[1036] The server decodes the generated tokenized data and creates a readable, formatted draft document. Based on the sentiment analysis results, the tone and content of the document are adjusted.

[1037] Input: Generated tokenized data

[1038] Output: Formatted draft document

[1039] Step 10:

[1040] The terminal displays the final draft document to the user. The user can review the generated document and make corrections or additions as needed.

[1041] Input: Formatted draft document

[1042] Output: Draft document displayed to the user

[1043] (Application Example 2)

[1044] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1045] Traditional customer service systems have struggled to analyze customer emotions in real time and respond accordingly. Furthermore, relying on manuals makes it difficult to improve customer satisfaction. A system is needed to solve these problems and provide more advanced customer service.

[1046] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1047] In this invention, the server includes means for collecting past document data, means for pre-training a natural language processing model using the past document data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating draft documents using the fine-tuning natural language processing model, means for providing the generated draft documents to the user, and means for analyzing the customer's emotions from their statements and facial expressions, generating a draft document based on the analysis results, and displaying it on a display device. This makes it possible to analyze the customer's emotions in real time and generate and provide an appropriate draft document based on the results.

[1048] "Past document data" refers to existing document information collected from within or outside the company, which is used for pre-training and analysis of the model.

[1049] A "natural language processing model" is a machine learning model designed to enable human understanding of natural language, and is a technology for automatically generating, analyzing, translating, and summarizing document content.

[1050] "Pre-training" is the process of using a broad dataset to teach a model basic language patterns and vocabulary before fine-tuning it for a specific task.

[1051] "Fine-tuning" is an additional learning process that uses specific parameters to further adapt a pre-trained model to a particular task or context.

[1052] A "user request" refers to the specific requests and parameters that a user provides to the system, and is the input information that the system uses to generate appropriate output based on these requests.

[1053] "Customer statements" refer to the words and sentences that customers make to the system, and by analyzing them, we can understand their intentions and requirements.

[1054] "Customer facial expressions" refer to the emotions and reactions shown by a customer's facial expressions, and by analyzing these, information is used to infer the customer's current emotional state.

[1055] "Emotional analysis" refers to the process of automatically identifying customer emotions from text and image data, and then generating appropriate responses and actions based on those results.

[1056] "Generating document drafts" means automatically creating appropriate document content using a natural language processing model based on user requests and analysis results.

[1057] "Displaying on a display device" means displaying information on a device that provides users with a visual representation of the generated document draft or analysis results, such as smart glasses or a display screen.

[1058] The present invention provides a system for pre-training a natural language processing model based on past document data, generating draft documents based on user requests and sentiment analysis results, and displaying them on a display device. The configuration and method for implementing this system are described below.

[1059] System Configuration

[1060] The system of the present invention consists of the following main elements:

[1061] 1. Server

[1062] Data collection method: Collect and store historical document data.

[1063] Natural language processing model: Pre-training is performed using collected document data, and then fine-tuning is performed using specific parameters (e.g., amount, timing).

[1064] Emotion analysis engine: Analyzes customer emotions from their statements and facial expressions.

[1065] 2. Terminal

[1066] Input method: Receives requests from users.

[1067] Display method: The generated document draft is displayed using smart glasses or a display.

[1068] Data collection and learning

[1069] The server collects historical document data from the company's database. This data is stored in text file or CSV format. The collected data is used to pre-train a natural language processing model (e.g., BERT). During pre-training, a tokenizer is used to tokenize the data, which is then input into the model. Next, the pre-trained model is fine-tuned with specific parameters.

[1070] Sentiment analysis and document generation

[1071] The user sends a request through a device (e.g., smart glasses). The request includes parameters such as "Payment amount: 5 million yen, Date: December 15, 2023". The device sends the user's speech and facial expressions to an emotion analysis engine to analyze the customer's emotions. Once the emotion analysis results are obtained, they are sent to the server.

[1072] The server generates a draft document using a finely tuned natural language processing model based on the received sentiment analysis results and request content. During this process, it adjusts the tone and content of the document, taking the sentiment analysis results into consideration. The generated draft document is then sent back to the terminal and displayed on the display device.

[1073] Hardware and software to use

[1074] 1. Hardware:

[1075] Smart glasses: Display customer information and analysis results in real time.

[1076] Webcam: Used to analyze customer facial expressions.

[1077] 2. Software:

[1078] Natural language processing model: BERT is used to analyze user utterances.

[1079] Sentiment analysis engine: Uses the TextBlob library to analyze emotions from customer statements.

[1080] OpenCV: Used to analyze customer facial expressions from image data.

[1081] Specific example

[1082] For example, if a customer says, "I'm having trouble with this product," the following process will occur:

[1083] 1. The user (store clerk) wears smart glasses and interacts with the customer.

[1084] 2. The user's spoken content and facial expression data captured by the webcam are sent to the emotion analysis engine, which then analyzes their emotions.

[1085] 3. The sentiment analysis results and request details are sent to the server, and a draft document is generated.

[1086] 4. The generated draft document is displayed on the smart glasses' screen.

[1087] Examples of prompts to input into a generative AI model include:

[1088] A customer asked a question about our company's products. They seemed confused, so please generate a reassuring message by providing FAQs and relevant product information.

[1089] This method allows users to obtain draft documents in real time that take into account appropriate and emotional tone, enabling them to provide customers with a more satisfying service.

[1090] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1091] Step 1:

[1092] The server collects historical document data from the company's database. This includes data in text file and CSV format. The collected document data is used as input to pre-train a natural language processing model.

[1093] Step 2:

[1094] The server tokenizes the collected document data using a tokenizer and inputs it into a natural language processing model (e.g., BERT). The model is then pre-trained to learn common document patterns and vocabulary. The output is the pre-trained model.

[1095] Step 3:

[1096] The server performs fine-tuning training on a pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning adapts the model to a specific business context, improving its accuracy. The output is the fine-tuned model.

[1097] Step 4:

[1098] The user sends a request from their device. The request includes the necessary parameters (e.g., "Payment amount: 5 million yen, Date: December 15, 2023"). The device sends the user's request to the sentiment analysis engine. The input is the user's request, and the output is the data transfer to the sentiment analysis engine.

[1099] Step 5:

[1100] The device analyzes the user's speech and facial expressions using an emotion analysis engine. The emotion analysis engine analyzes emotions from input text and image data and sends the results to the server. The input is the user's speech and facial expression data, and the output is the emotion analysis result.

[1101] Step 6:

[1102] The server inputs the generated document draft, based on the sentiment analysis results and the user's request, into a finely tuned natural language processing model. The model generates the document draft taking the sentiment analysis results into account. The input is the sentiment analysis results and the user's request, and the output is the generated document draft.

[1103] Step 7:

[1104] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Based on the sentiment analysis results, the tone and content of the document are adjusted. The input is the generated tokenized data, and the output is the draft document in an easy-to-read format.

[1105] Step 8:

[1106] The terminal displays the generated document draft on a display device (e.g., smart glasses or a screen). This allows the user to quickly obtain a document draft that is appropriate and considers emotional tone. The input is a readable document draft, and the output is a display on a display device.

[1107] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1108] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1109] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1110] [Fourth Embodiment]

[1111] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1112] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1113] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1114] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1115] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1116] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1117] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1118] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1119] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1120] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1121] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1122] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1123] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1124] The system of the present invention automatically generates draft documents based on user requests by collecting historical document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. This system is implemented according to the following procedure.

[1125] Program processing

[1126] 1. The server first collects historical document data from the company's internal database. This includes various official documents, such as approval forms. The collected data is saved in text file or CSV format.

[1127] 2. The server then sets up a natural language processing model (e.g., BERT). It loads the collected historical document data and uses this data to pre-train the model. It tokenizes the document data using a tokenizer and feeds it to the model as input to learn patterns and vocabulary.

[1128] 3. The server performs fine-tuning learning on the pre-trained model using specific parameters (e.g., amount or timing). This allows the model to adapt to a specific business context and generate more specific and accurate document drafts. Fine-tuning learning is also performed by tokenizing data using a tokenizer and feeding it to the model.

[1129] 4. The user sends a request to the chatbot to create an approval document using their own device. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[1130] 5. The terminal sends the user's request to the server. The server receives this request, tokenizes the request content, and provides it to the refined model.

[1131] 6. The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Specifically, the content generated by the model is processed into a draft document and converted into a format that the user can review.

[1132] 7. The terminal displays the generated document draft to the user. This allows the user to obtain a suitable document draft in a short amount of time.

[1133] Specific example

[1134] For example, if a user sends a request to the chatbot to create an approval document for a payment of 5 million yen on December 15, 2023, the following process will occur:

[1135] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[1136] 2. The device sends this request to the server.

[1137] 3. The server tokenizes the request and feeds it to a finely tuned natural language processing model.

[1138] 4. The server generates a draft approval document and decodes it as follows: "Proposal: The budget for this project is 5 million yen, and the implementation date is December 15, 2023."

[1139] 5. The terminal displays this draft document to the user.

[1140] In this way, the system of the present invention provides an environment in which users can easily create approval documents, thereby reducing time and effort.

[1141] The following describes the processing flow.

[1142] Step 1:

[1143] The server collects historical document data from the company's internal database. This includes data from official documents such as approval requests and reports. The collected data is stored in a single text file or a CSV file.

[1144] Step 2:

[1145] The server initializes a natural language processing model (e.g., BERT). This involves using a tool called a tokenizer to tokenize the collected text data (divide it into units of words or sentences) and convert it into a format that the model can understand.

[1146] Step 3:

[1147] The server inputs tokenized data into a natural language processing model for pre-training. This allows the model to learn common document patterns and vocabulary from past document data.

[1148] Step 4:

[1149] The server prepares specific parameters (e.g., amount and time period). Based on these specific parameters, it further refines the pre-trained model. This improves the model's accuracy and adaptability to the specific business context.

[1150] Step 5:

[1151] The user sends a request to the chatbot from their device to create an approval document. This request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[1152] Step 6:

[1153] The terminal sends data to the server via a communication interface in order to send user requests to the server.

[1154] Step 7:

[1155] The server tokenizes the user's request using a tokenizer. This tokenized data is then input into a finely tuned natural language processing model.

[1156] Step 8:

[1157] The server decodes the tokenized data generated from the model and creates a draft document in an easy-to-read format. Specifically, it converts the generated tokenized data into natural language sentences.

[1158] Step 9:

[1159] The terminal displays the draft document sent from the server to the user. This allows the user to quickly review an appropriate draft document.

[1160] This series of processes allows users to efficiently create approval documents, significantly reducing the time and effort required.

[1161] (Example 1)

[1162] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1163] In modern businesses, document creation is a time-consuming and labor-intensive task. In particular, creating new documents based on past documents is cumbersome, requiring careful reference and editing, and maintaining consistency is difficult. Furthermore, generating draft documents quickly and accurately in response to specific user requests is challenging to do manually. This leads to a decrease in overall business efficiency.

[1164] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1165] In this invention, the server includes means for collecting historical document data, means for pre-training a natural language processing model using the historical document data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating document drafts using the fine-tuning natural language processing model, means for providing the generated document drafts to users, means for tokenizing the content of the user's request, means for inputting the tokenized data into the natural language processing model, and means for decoding the tokenized data generated from the natural language processing model. This makes it possible for users to easily and efficiently create high-quality documents.

[1166] "Past document data" refers to information about previously created documents extracted from a company's internal document database.

[1167] A "natural language processing model" is a machine learning model that analyzes and understands text data, and in particular, it uses deep learning techniques to learn the meaning of sentences.

[1168] "Pre-training" is the process of training a natural language processing model using a large amount of text data to learn basic language patterns and structures.

[1169] "Fine-tuning learning" is the process of setting specific parameters for a pre-trained natural language processing model and making adjustments to suit a more specific context.

[1170] "Specific parameters" refer to values ​​or conditions (e.g., amount, time) that are important elements when generating a document.

[1171] A "user request" refers to instructions or requests regarding document generation that a user sends to the system via their device.

[1172] "Tokenization" is the process of dividing text data into smaller units (tokens) such as words and phrases.

[1173] "Decoding" is the process of converting tokenized data back into natural language text in a format that is easy for humans to read.

[1174] A "draft document" is a preliminary version of a document generated by a natural language processing model and created based on user requests.

[1175] This invention is a system for efficiently creating documents within a company. It collects past document data, pre-trains a natural language processing model using that data, and then performs fine-tuning learning using specific parameters, thereby automatically generating document drafts based on user requests.

[1176] First, the server collects historical document data from the company's internal database. This data includes various types of documents such as approval forms, reports, and contracts. The collected data is saved in text or CSV format in preparation for the next processing step.

[1177] Next, the server prepares a natural language processing model (e.g., BERT) using the Python Torch library. It loads historical document data and uses this data to pre-train the model. Specifically, it tokenizes the document data using a tokenizer (e.g., BertTokenizer) and inputs this tokenized data into the model. This allows the model to learn document patterns and vocabulary.

[1178] Once pre-training is complete, the server fine-tunes the model using specific parameters (e.g., amount or time period). This fine-tuning is also performed by tokenizing data using a tokenizer and feeding it to the model. This allows the model to generate more specific and accurate document drafts adapted to the particular business context.

[1179] The user sends a request to the chatbot to create an approval document using their device. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023"). The device then sends this request to the server.

[1180] The server tokenizes the received request and feeds it to a finely tuned model. The model generates a draft document based on the request. The generated tokenized data is decoded again into a readable draft document. Finally, the terminal displays this draft document to the user.

[1181] Specific example

[1182] For example, if a user sends a request to the chatbot to create an approval document for a payment of 5 million yen on December 15, 2023, the following process will occur:

[1183] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[1184] 2. The device sends this request to the server.

[1185] 3. The server tokenizes the request and feeds it to a finely tuned natural language processing model.

[1186] 4. The server generates a draft approval document using a natural language processing model.

[1187] 5. The server decodes the generated content and creates a draft document stating, "Proposed content: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023."

[1188] 6. The terminal displays this draft document to the user.

[1189] This system allows users to obtain appropriate and consistent document drafts in a short amount of time, thereby improving work efficiency.

[1190] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1191] Step 1:

[1192] The server executes a mechanism to collect historical document data from the company's internal database. It uses database queries to retrieve document data in text or CSV format, such as approval forms, reports, and contracts. The database queries used as input output document data, which is saved to the "past_documents.csv" file. Specifically, it calls an API for database access and executes SQL queries.

[1193] Step 2:

[1194] The server prepares and configures a natural language processing model (e.g., BERT). It loads the model using the Python Torch library and sets the tokenizer, for example, `tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')`. It reads a past document data file, "past_documents.csv", as input and tokenizes it using the tokenizer. The output is the tokenized document data.

[1195] Step 3:

[1196] The server pre-trains a BERT model using past document data. The tokenized data from the tokenizer is created as a dataset using `dataset = TextDataset(tokenizer, 'past_documents.csv')`, and the model is pre-trained using `model.train(dataset)`. The pre-trained model is then output using the tokenized data output from the tokenizer as input.

[1197] Step 4:

[1198] The server performs fine-tuning on a pre-trained model using specific parameters (e.g., amount or time period). It sets specific parameters, tokenizes the fine-tuning data using a tokenizer, and feeds it to the model. Specifically, it executes `model.fine_tune(tokenized_data, parameters)`. The input is the specific parameters and tokenized fine-tuning data, and the output is the fine-tuned model.

[1199] Step 5:

[1200] The user sends a request for the creation of an approval document to the chatbot via their device. Specifically, they enter "Amount: 5 million yen, Date: December 15, 2023" into the chatbot's interface and press the send button. The input is the user's prompt text, and the output is the request sent by the chatbot's interface.

[1201] Step 6:

[1202] The terminal sends the user's request to the server. Specifically, it creates an HTTP POST request and sends the request data to a specific endpoint on the server (e.g., " / generate_document"). The input is the user's request, and the output is the sending of the HTTP request to the server.

[1203] Step 7:

[1204] The server tokenizes the request content and provides it to a fine-tuned model. The received request is tokenized using `tokenizer.tokenize(request_data)`, and a draft document is generated using `model.generate(tokenized_request)`. The input is the tokenized request data, and the output is the generated draft document.

[1205] Step 8:

[1206] The server decodes the generated tokenized data and creates a draft document. The token output is decoded using tokenizer.decode(generated_tokens) and converted into a user-readable format. The input is the tokenized draft document data, and the output is the decoded draft document.

[1207] Step 9:

[1208] The terminal displays the generated document draft to the user. Specifically, it receives the document draft from the server and displays it on the chatbot screen. The input is the document draft from the server, and the output is the display to the user.

[1209] (Application Example 1)

[1210] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1211] Logistics centers require a wide variety of documents (purchase orders, shipping instructions, inventory management reports, etc.), but creating these documents is time-consuming and labor-intensive. Furthermore, because these documents must accurately reflect past data and current conditions, they are prone to human error. While efficiency is needed in these document creation processes, on-site personnel often require advanced specialized knowledge. A system is needed to address these challenges.

[1212] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1213] In this invention, the server includes means for collecting historical data, means for pre-training a natural language processing model using the historical data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating draft documents using the fine-tuning natural language processing model, means for providing the generated draft documents to users, means for collecting historical order history and inventory data from the logistics center's database, means for automatically generating purchase orders, shipping instructions, and inventory management reports using the natural language processing model, and means for displaying and modifying the generated draft documents on a smartphone. This makes it possible to automatically generate documents efficiently and accurately at the logistics center, significantly reducing the burden on on-site staff and improving work efficiency.

[1214] "Past data" refers to information including past order history and inventory data at the logistics center.

[1215] A "natural language processing model" is a machine learning model that understands human language and generates documents.

[1216] "Pre-training" is the process of training a natural language processing model from the ground up using a large amount of data.

[1217] "Fine-tuning learning" is the process of optimizing a natural language processing model, which has been pre-trained using specific parameters, to suit a particular application.

[1218] A "user request" is information that a user uses to request document generation based on specific parameters.

[1219] A "draft document" is a document generated by a natural language processing model and submitted as a candidate.

[1220] "Means of provision" refers to the method by which the generated draft document is displayed for the user to review and modify.

[1221] A "logistics center database" is a data system that stores order history and inventory data managed within a logistics center.

[1222] A "purchase order" is a document that details the order for goods or services.

[1223] A "shipping instruction sheet" is a document that contains instructions for shipping goods out of a logistics center.

[1224] An "inventory management report" is a report that summarizes the current inventory status and related information.

[1225] A "smartphone" is a portable, multi-functional device that can display and process information using applications.

[1226] This invention is a system for efficiently and automatically generating documents in a logistics center. This system collects historical data, uses that data to pre-train a natural language processing model, and then fine-tunes it using specific parameters. Next, it generates draft documents based on user requests and provides them to the user.

[1227] The specific configuration of this system is as follows:

[1228] Hardware and software:

[1229] Server: Responsible for data collection, model training, fine-tuning, and document generation. The natural language processing model used includes BERT.

[1230] Database system: Stores order history and inventory data for the logistics center. Uses MySQL, PostgreSQL, etc.

[1231] Tokenizer: Uses a Python library (such as Hugging Face's Transformers) to tokenize the collected data.

[1232] User interface device (smartphone): Built with React Native to display generated document drafts and receive user requests.

[1233] Program processing:

[1234] 1. Data collection:

[1235] The server collects past order history and inventory data from the logistics center's database. The data is saved in text file or CSV format.

[1236] 2. Pre-learning:

[1237] The server sets up a natural language processing model (BERT) and pre-trains the model using collected historical data. The data is tokenized using the Python Hugging Face library and input into the model.

[1238] 3. Fine-tuning learning:

[1239] The server fine-tunes a pre-trained model using specific parameters (product name, quantity, delivery date). This process also uses a tokenizer to tokenize the data and add optimization information to the model.

[1240] 4. Receiving and sending user requests:

[1241] Users use their smartphones to submit requests for new purchase orders or shipping orders. These requests include the necessary parameters (e.g., "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023").

[1242] The smartphone sends this request to the server.

[1243] 5. Document draft generation:

[1244] The server tokenizes the received request and feeds it into a finely tuned model. The model generates a draft document based on the specified parameters.

[1245] 6. Providing draft documents:

[1246] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. For example, it might generate a specific draft document such as, "Please deliver 100 units of product A on December 1, 2023."

[1247] The generated document draft is displayed to the user on their smartphone. The user can review the content and make corrections as needed.

[1248] This will enable the automation and streamlining of document generation tasks in logistics centers. Furthermore, it will allow users to obtain accurate documents in a shorter time, improving both work efficiency and accuracy.

[1249] Specific example:

[1250] When a user submits a request to create a purchase order with the following information: "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023," the server tokenizes this information and feeds it into a finely tuned natural language processing model to generate a purchase order like this: "Please deliver 100 units of Product A on December 1, 2023." This draft document is then displayed to the user on their smartphone.

[1251] Example of a prompt:

[1252] "Please create a purchase order using the following parameters: Product Name: Product A, Quantity: 100, Due Date: December 1, 2023"

[1253] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1254] Step 1:

[1255] The server collects past order history and inventory data from the logistics center's database. It queries the database system (e.g., MySQL or PostgreSQL) to extract the necessary data. The collected data is saved in text file or CSV format.

[1256] Input: Logistics center database

[1257] Output: Text or CSV file containing past order history and inventory data.

[1258] Step 2:

[1259] The server pre-trains a natural language processing model (BERT) using the collected historical data. The data is tokenized using the Python Hugging Face library and input into the model. This allows the model to learn basic language structures and patterns.

[1260] Input: Text or CSV file containing past order history and inventory data.

[1261] Output: Pre-trained natural language processing model

[1262] Step 3:

[1263] The server fine-tunes a pre-trained model using specific parameters (product name, quantity, delivery date). The Python Hugging Face library is used to tokenize additional data and add optimization information to the model. This allows the model to adapt to specific business contexts.

[1264] Input: Specific parameters (product name, quantity, delivery date)

[1265] Output: Fine-tuned natural language processing model

[1266] Step 4:

[1267] Users use their smartphones to submit requests for new purchase orders or shipping orders. These requests include the necessary parameters (e.g., "Product Name: Product A, Quantity: 100, Delivery Date: December 1, 2023").

[1268] Input: Request parameters (product name, quantity, delivery date)

[1269] Output: Request from the user's smartphone

[1270] Step 5:

[1271] The terminal sends the user's request to the server. The request data is transferred to the server using protocols such as HTTP requests.

[1272] Input: Request from the user's smartphone

[1273] Output: Request data to the server

[1274] Step 6:

[1275] The server tokenizes the received request using a tokenizer and feeds it into a finely tuned model. The model generates a draft document based on the specified parameters.

[1276] Input: Request data to the server

[1277] Output: Tokenized data of the generated document

[1278] Step 7:

[1279] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Specifically, it converts the tokenized data into natural language sentences and formats it as a structured document.

[1280] Input: Tokenized data of the generated document

[1281] Output: Draft document in an easy-to-read format

[1282] Step 8:

[1283] The terminal displays the generated document draft to the user. The user can review the document draft through the smartphone interface and make revisions as needed.

[1284] Input: Draft document in an easy-to-read format

[1285] Output: Draft document that users can review and edit.

[1286] The above steps automate and streamline document generation at the logistics center. Clearly defining the specific actions and data input / output performed at each step makes system implementation and operation easier.

[1287] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1288] The system of the present invention automatically generates draft documents based on user requests by collecting past document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the content and tone of the draft documents can be adjusted according to the user's emotions. This system is implemented according to the following procedure.

[1289] Program processing

[1290] 1. The server collects historical document data from the company's internal database. The collected data is saved in text file or CSV format.

[1291] 2. The server sets up a natural language processing model (e.g., BERT). The collected document data is tokenized using a tokenizer and input into the model. The model is then pre-trained to learn common document patterns and vocabulary.

[1292] 3. The server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning allows the model to adapt to the specific business context and improves its accuracy.

[1293] 4. The user sends a request to the chatbot from their device to create a proposal document. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[1294] 5. The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes emotions from the input text or voice and sends the results to the server.

[1295] 6. The server tokenizes the user's request based on the sentiment analysis results received from the sentiment engine and inputs it into a finely tuned model. The model generates a draft document taking the sentiment analysis results into account.

[1296] 7. The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Based on the sentiment analysis results, the tone and content of the document are adjusted.

[1297] 8. The terminal displays the generated draft document to the user. This allows the user to quickly obtain a draft document that takes into account an appropriate and emotionally balanced tone.

[1298] Specific example

[1299] For example, if a user submits a request to create a proposal document for a transaction with the following conditions: "Amount: 5 million yen, Date: December 15, 2023, Emotion: Angry," the following process will occur:

[1300] 1. User: Please prepare a proposal for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[1301] 2. The device sends this request to the emotion engine to analyze the user's emotions. If the analysis result is "anger," the tone is taken into consideration.

[1302] 3. The device sends the request to the server along with the sentiment analysis results.

[1303] 4. The server tokenizes the request and feeds it into a finely tuned natural language processing model. The model generates a draft document, taking sentiment analysis results into account.

[1304] 5. The server decodes the draft document and generates content such as, "Proposal: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[1305] 6. The terminal displays this draft document to the user.

[1306] This series of processes allows users to efficiently create approval documents that take emotions into consideration, thereby improving the accuracy and effectiveness of communication.

[1307] The following describes the processing flow.

[1308] Step 1:

[1309] The server collects historical document data from the company's internal database. The collected data is saved in text file or CSV format.

[1310] Step 2:

[1311] The server configures a natural language processing model (e.g., BERT). It tokenizes the collected document data using a tokenizer. The tokenizer divides the document data into units of words and phrases, converting it into a format that the model can understand.

[1312] Step 3:

[1313] The server inputs tokenized data into a natural language processing model for pre-training. The model learns common document patterns and vocabulary from past document data.

[1314] Step 4:

[1315] The server performs fine-tuning learning on a pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning allows the model to adapt to a specific business context, improving its accuracy and adaptability.

[1316] Step 5:

[1317] The user sends a request to the chatbot from their device to create an approval document. The request includes the necessary parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[1318] Step 6:

[1319] The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the emotions from the input text or voice and sends the results to the server.

[1320] Step 7:

[1321] The server tokenizes the user's request based on the sentiment analysis results received from the sentiment engine and inputs it into a finely tuned model. The model then generates a draft document, taking the sentiment analysis results into account.

[1322] Step 8:

[1323] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. The generated draft document is then adjusted in tone and content based on the sentiment analysis results.

[1324] Step 9:

[1325] The terminal displays the generated draft document to the user. This allows the user to quickly review a draft document that takes into account appropriate and emotional tone.

[1326] Specific example

[1327] For example, if a user submits a request to create a proposal document for a transaction with the following conditions: "Amount: 5 million yen, Date: December 15, 2023, Emotion: Angry," the following process will occur:

[1328] Step 1:

[1329] The user submits a request from their device to create an approval document for a transaction with the following details: "Payment amount: 5 million yen, Date: December 15, 2023".

[1330] Step 2:

[1331] The device sends this request to the emotion engine, which analyzes the user's emotions. The analysis results detect that the user is "angry."

[1332] Step 3:

[1333] The emotion engine sends the emotion analysis results to the server.

[1334] Step 4:

[1335] The server tokenizes the request along with the sentiment analysis results and inputs them into a finely tuned model. The model then generates a draft document, taking the sentiment analysis results into account.

[1336] Step 5:

[1337] The server decodes the generated draft document and produces content such as, "Proposal: The budget for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[1338] Step 6:

[1339] The terminal displays the generated draft document to the user.

[1340] In this way, users can create emotionally conscious document drafts quickly and efficiently.

[1341] (Example 2)

[1342] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1343] Conventional document generation systems failed to consider user emotions, resulting in a uniform tone in the generated documents. This led to the problem of documents not adequately reflecting user intent and feelings. Furthermore, methods for providing highly accurate models adapted to business contexts through fine-tuning learning using past document data were insufficient.

[1344] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1345] In this invention, the server includes means for collecting past document data, means for pre-training a natural language processing model using the past document data, means for fine-tuning the natural language processing model using specific parameters, means for analyzing user requests to recognize emotional states, and means for adjusting the tone and content of draft documents considering the emotional states. This makes it possible to generate documents with an appropriate tone that reflects the user's emotions. Furthermore, it is possible to provide highly accurate documents adapted to a specific business context.

[1346] "Past document data" refers to documents previously created by companies or individuals that are stored in databases or file systems.

[1347] A "natural language processing model" is an artificial intelligence model used to understand and generate human language, and includes models such as BERT.

[1348] "Pre-training" is the process of initially training a model using a large amount of collected data to teach it common language patterns and vocabulary.

[1349] "Fine-tuning learning" is the process of further training a pre-trained model according to specific tasks or contexts to improve the model's accuracy.

[1350] A "user request" is input from a user to give specific instructions or requests to the system, and is sent through a chatbot or form.

[1351] A "document draft" is a draft of a document, such as a proposal or report, generated through user requests and system processing.

[1352] "Emotional state" refers to the state of emotions analyzed from the user's input, and includes emotions such as anger, joy, and sadness.

[1353] "Tokenization" is the process of breaking down text data into units of words or subwords and assigning an ID to each of them.

[1354] An "emotion engine" is software or an algorithm that analyzes a user's text or voice input to recognize their emotional state.

[1355] "Decoding" is the process of converting tokenized data back into a human-readable text format.

[1356] The system of this invention automatically generates draft documents based on user requests by collecting past document data, pre-training a natural language processing model using that data, and then performing fine-tuning learning using specific parameters. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, the content and tone of the draft documents can be adjusted according to the user's emotions.

[1357] Specifically, the server first collects historical document data from the company's internal database. This data includes business documents such as sales reports, meeting minutes, and approval documents. The collected data is saved in text file or CSV format.

[1358] Next, the server loads a natural language processing model (e.g., BERT) using the Hugging Face library. The collected document data is tokenized using a tokenizer (e.g., BERT Tokenizer) and input into the model. The model is pre-trained using this tokenized data to learn common language patterns and vocabulary.

[1359] Subsequently, the server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount and timing). This fine-tuning allows the model to adapt to the specific business context, enabling more accurate document generation. The fine-tuned model is stored on the server and used to respond to user requests.

[1360] The user uses a chatbot interface from their device to request the creation of an approval document for specific details (e.g., "Amount: 5 million yen, Date: December 15, 2023"). This request is sent from the device to the emotion engine, which analyzes the user's emotional state based on their input. The emotion engine classifies the user's emotions into categories such as "anger," "joy," and "sadness," and returns the result to the device.

[1361] The terminal sends the sentiment analysis results and the user's request to the server, which tokenizes the request and inputs it into a finely tuned model. The model takes the sentiment analysis results into account and generates a draft document that reflects the user's emotions. For example, if anger is detected, the tone of the document is strengthened and adjusted to convey a request for immediate action.

[1362] The generated tokenized data is decoded on the server and sent to the terminal as a readable, formatted draft document. The terminal displays the final draft document to the user, who then reviews the generated document and makes any necessary corrections or additions.

[1363] As a concrete example, consider a case where a user submits a request to create a proposal document with the following conditions: "Payment amount: 5 million yen, Date: December 15, 2023, Emotion: Angry." In response to this request, the server generates a draft document reflecting the user's emotion and provides it to the user. The generated draft document would be displayed with content such as: "Proposal: The payment amount for this project will be 5 million yen, and the implementation date will be December 15, 2023. This decision is urgently needed."

[1364] This system allows users to efficiently create documents that take emotions into account, thereby improving the accuracy and effectiveness of communication.

[1365] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1366] Step 1:

[1367] The server collects historical document data from the company's internal database. This includes business documents such as sales reports, meeting minutes, and approval documents. The collected data is saved in text file or CSV format.

[1368] Input: Corporate database

[1369] Output: Collected historical document data (text files, CSV)

[1370] Step 2:

[1371] The server loads a natural language processing model (e.g., BERT) using the Hugging Face library. It tokenizes the collected document data using a tokenizer (e.g., BERT Tokenizer). Each document is broken down into words and subwords, and converted into integer IDs.

[1372] Input: Past document data

[1373] Output: Tokenized document data

[1374] Step 3:

[1375] The server inputs tokenized data into the model and performs pre-training. During this process, the model learns document patterns and vocabulary. An initially trained model is obtained.

[1376] Input: Tokenized document data

[1377] Output: Pre-trained natural language processing model

[1378] Step 4:

[1379] The server performs fine-tuning training on the pre-trained model using specific parameters (e.g., amount or timing). This allows the model to adapt to the specific business context and improve its accuracy.

[1380] Input: Pre-trained natural language processing model, specific parameters

[1381] Output: Fine-tuned natural language processing model

[1382] Step 5:

[1383] The user sends a request for the creation of an approval document using the chatbot interface from their device. The request includes specific parameters (e.g., "Amount: 5 million yen, Date: December 15, 2023").

[1384] Input: User request (payment amount, timing, etc.)

[1385] Output: Request data received by the terminal

[1386] Step 6:

[1387] The device sends the user's request to the emotion engine, which analyzes the user's emotions. The emotion engine classifies the emotional state from the input text or voice and returns the result to the device.

[1388] Input: User request, emotion engine

[1389] Output: Emotion analysis results (anger, joy, sadness, etc.)

[1390] Step 7:

[1391] The device sends the sentiment analysis results and the user's request to the server.

[1392] Input: Sentiment analysis results, user requests

[1393] Output: Sentiment analysis results and requests received by the server

[1394] Step 8:

[1395] The server tokenizes the request and inputs it into a finely tuned model. The model takes sentiment analysis results into account and generates a draft document that reflects the user's emotions.

[1396] Input: Tokenized request content, sentiment analysis results, and a fine-tuned natural language processing model.

[1397] Output: Generated draft document

[1398] Step 9:

[1399] The server decodes the generated tokenized data and creates a readable, formatted draft document. Based on the sentiment analysis results, the tone and content of the document are adjusted.

[1400] Input: Generated tokenized data

[1401] Output: Formatted draft document

[1402] Step 10:

[1403] The terminal displays the final draft document to the user. The user can review the generated document and make corrections or additions as needed.

[1404] Input: Formatted draft document

[1405] Output: Draft document displayed to the user

[1406] (Application Example 2)

[1407] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1408] Traditional customer service systems have struggled to analyze customer emotions in real time and respond accordingly. Furthermore, relying on manuals makes it difficult to improve customer satisfaction. A system is needed to solve these problems and provide more advanced customer service.

[1409] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1410] In this invention, the server includes means for collecting past document data, means for pre-training a natural language processing model using the past document data, means for fine-tuning the natural language processing model using specific parameters, means for receiving requests from users and generating draft documents using the fine-tuning natural language processing model, means for providing the generated draft documents to the user, and means for analyzing the customer's emotions from their statements and facial expressions, generating a draft document based on the analysis results, and displaying it on a display device. This makes it possible to analyze the customer's emotions in real time and generate and provide an appropriate draft document based on the results.

[1411] "Past document data" refers to existing document information collected from within or outside the company, which is used for pre-training and analysis of the model.

[1412] A "natural language processing model" is a machine learning model designed to enable human understanding of natural language, and is a technology for automatically generating, analyzing, translating, and summarizing document content.

[1413] "Pre-training" is the process of using a broad dataset to teach a model basic language patterns and vocabulary before fine-tuning it for a specific task.

[1414] "Fine-tuning" is an additional learning process that uses specific parameters to further adapt a pre-trained model to a particular task or context.

[1415] A "user request" refers to the specific requests and parameters that a user provides to the system, and is the input information that the system uses to generate appropriate output based on these requests.

[1416] "Customer statements" refer to the words and sentences that customers make to the system, and by analyzing them, we can understand their intentions and requirements.

[1417] "Customer facial expressions" refer to the emotions and reactions shown by a customer's facial expressions, and by analyzing these, information is used to infer the customer's current emotional state.

[1418] "Emotional analysis" refers to the process of automatically identifying customer emotions from text and image data, and then generating appropriate responses and actions based on those results.

[1419] "Generating document drafts" means automatically creating appropriate document content using a natural language processing model based on user requests and analysis results.

[1420] "Displaying on a display device" means displaying information on a device that provides users with a visual representation of the generated document draft or analysis results, such as smart glasses or a display screen.

[1421] The present invention provides a system for pre-training a natural language processing model based on past document data, generating draft documents based on user requests and sentiment analysis results, and displaying them on a display device. The configuration and method for implementing this system are described below.

[1422] System Configuration

[1423] The system of the present invention consists of the following main elements:

[1424] 1. Server

[1425] Data collection method: Collect and store historical document data.

[1426] Natural language processing model: Pre-training is performed using collected document data, and then fine-tuning is performed using specific parameters (e.g., amount, timing).

[1427] Emotion analysis engine: Analyzes customer emotions from their statements and facial expressions.

[1428] 2. Terminal

[1429] Input method: Receives requests from users.

[1430] Display method: The generated document draft is displayed using smart glasses or a display.

[1431] Data collection and learning

[1432] The server collects historical document data from the company's database. This data is stored in text file or CSV format. The collected data is used to pre-train a natural language processing model (e.g., BERT). During pre-training, a tokenizer is used to tokenize the data, which is then input into the model. Next, the pre-trained model is fine-tuned with specific parameters.

[1433] Sentiment analysis and document generation

[1434] The user sends a request through a device (e.g., smart glasses). The request includes parameters such as "Payment amount: 5 million yen, Date: December 15, 2023". The device sends the user's speech and facial expressions to an emotion analysis engine to analyze the customer's emotions. Once the emotion analysis results are obtained, they are sent to the server.

[1435] The server generates a draft document using a finely tuned natural language processing model based on the received sentiment analysis results and request content. During this process, it adjusts the tone and content of the document, taking the sentiment analysis results into consideration. The generated draft document is then sent back to the terminal and displayed on the display device.

[1436] Hardware and software to use

[1437] 1. Hardware:

[1438] Smart glasses: Display customer information and analysis results in real time.

[1439] Webcam: Used to analyze customer facial expressions.

[1440] 2. Software:

[1441] Natural language processing model: BERT is used to analyze user utterances.

[1442] Sentiment analysis engine: Uses the TextBlob library to analyze emotions from customer statements.

[1443] OpenCV: Used to analyze customer facial expressions from image data.

[1444] Specific example

[1445] For example, if a customer says, "I'm having trouble with this product," the following process will occur:

[1446] 1. The user (store clerk) wears smart glasses and interacts with the customer.

[1447] 2. The user's spoken content and facial expression data captured by the webcam are sent to the emotion analysis engine, which then analyzes their emotions.

[1448] 3. The sentiment analysis results and request details are sent to the server, and a draft document is generated.

[1449] 4. The generated draft document is displayed on the smart glasses' screen.

[1450] Examples of prompts to input into a generative AI model include:

[1451] A customer asked a question about our company's products. They seemed confused, so please generate a reassuring message by providing FAQs and relevant product information.

[1452] This method allows users to obtain draft documents in real time that take into account appropriate and emotional tone, enabling them to provide customers with a more satisfying service.

[1453] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1454] Step 1:

[1455] The server collects historical document data from the company's database. This includes data in text file and CSV format. The collected document data is used as input to pre-train a natural language processing model.

[1456] Step 2:

[1457] The server tokenizes the collected document data using a tokenizer and inputs it into a natural language processing model (e.g., BERT). The model is then pre-trained to learn common document patterns and vocabulary. The output is the pre-trained model.

[1458] Step 3:

[1459] The server performs fine-tuning training on a pre-trained model using specific parameters (e.g., amount or timing). This fine-tuning adapts the model to a specific business context, improving its accuracy. The output is the fine-tuned model.

[1460] Step 4:

[1461] The user sends a request from their device. The request includes the necessary parameters (e.g., "Payment amount: 5 million yen, Date: December 15, 2023"). The device sends the user's request to the sentiment analysis engine. The input is the user's request, and the output is the data transfer to the sentiment analysis engine.

[1462] Step 5:

[1463] The device analyzes the user's speech and facial expressions using an emotion analysis engine. The emotion analysis engine analyzes emotions from input text and image data and sends the results to the server. The input is the user's speech and facial expression data, and the output is the emotion analysis result.

[1464] Step 6:

[1465] The server inputs the generated document draft, based on the sentiment analysis results and the user's request, into a finely tuned natural language processing model. The model generates the document draft taking the sentiment analysis results into account. The input is the sentiment analysis results and the user's request, and the output is the generated document draft.

[1466] Step 7:

[1467] The server decodes the generated tokenized data and creates a draft document in an easy-to-read format. Based on the sentiment analysis results, the tone and content of the document are adjusted. The input is the generated tokenized data, and the output is the draft document in an easy-to-read format.

[1468] Step 8:

[1469] The terminal displays the generated document draft on a display device (e.g., smart glasses or a screen). This allows the user to quickly obtain a document draft that is appropriate and considers emotional tone. The input is a readable document draft, and the output is a display on a display device.

[1470] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1471] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1472] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1473] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1474] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1475] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1476] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1477] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1478] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1479] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1480] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1481] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1482] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1483] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1484] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1485] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1486] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1487] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1488] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1489] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1490] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1491] The following is further disclosed regarding the embodiments described above.

[1492] (Claim 1)

[1493] Means for collecting historical document data,

[1494] A means for pre-training a natural language processing model using the aforementioned past document data,

[1495] A means for fine-tuning the natural language processing model using specific parameters,

[1496] A means for receiving a request from a user and generating a draft document using the finely tuned natural language processing model,

[1497] Means for providing the generated draft document to the user,

[1498] A system that includes this.

[1499] (Claim 2)

[1500] The specific parameters for the aforementioned fine-tuning learning include the amount and the timing.

[1501] The system according to claim 1.

[1502] (Claim 3)

[1503] The historical document data was extracted from the company's internal document database.

[1504] The system according to claim 1.

[1505] "Example 1"

[1506] (Claim 1)

[1507] Means for collecting historical document data,

[1508] A means for pre-training a natural language processing model using the aforementioned past document data,

[1509] A means for fine-tuning the natural language processing model using specific parameters,

[1510] A means for receiving a request from a user and generating a draft document using the finely tuned natural language processing model,

[1511] Means for providing the generated draft document to the user,

[1512] A means of tokenizing the user's request content,

[1513] A means of inputting tokenized data into a natural language processing model,

[1514] A means for decoding tokenized data generated from the aforementioned natural language processing model,

[1515] A system that includes this.

[1516] (Claim 2)

[1517] The specific parameters for the aforementioned fine-tuning learning include the amount and the timing.

[1518] The system according to claim 1.

[1519] (Claim 3)

[1520] The historical document data was extracted from the company's internal document database.

[1521] The system according to claim 1.

[1522] "Application Example 1"

[1523] (Claim 1)

[1524] Means of collecting past data,

[1525] A means for pre-training a natural language processing model using the aforementioned past data,

[1526] A means for fine-tuning the natural language processing model using specific parameters,

[1527] A means for receiving a request from a user and generating a draft document using the finely tuned natural language processing model,

[1528] Means for providing the generated draft document to the user,

[1529] A method for collecting past order history and inventory data from the logistics center's database,

[1530] A means for automatically generating purchase orders, shipping instructions, and inventory management reports using the aforementioned natural language processing model,

[1531] A means for displaying and modifying the generated draft document on a smartphone,

[1532] A system that includes this.

[1533] (Claim 2)

[1534] The specific parameters for the aforementioned fine-tuning learning include product name, quantity, and delivery date.

[1535] The system according to claim 1.

[1536] (Claim 3)

[1537] The historical data was extracted from the logistics center's database.

[1538] The system according to claim 1.

[1539] "Example 2 of combining an emotion engine"

[1540] (Claim 1)

[1541] Means for collecting historical document data,

[1542] A means for pre-training a natural language processing model using the aforementioned past document data,

[1543] A means for fine-tuning the natural language processing model using specific parameters,

[1544] A means for receiving a request from a user and generating a draft document using the finely tuned natural language processing model,

[1545] Means for providing the generated draft document to the user,

[1546] A means of analyzing user requests to recognize their emotional state,

[1547] Means for adjusting the tone and content of the draft document in consideration of the aforementioned emotional state,

[1548] A system that includes this.

[1549] (Claim 2)

[1550] The system according to claim 1, wherein the specific parameters for the fine-tuning learning include an amount and a time.

[1551] (Claim 3)

[1552] The system according to claim 1, wherein past document data is extracted from a document database within the company.

[1553] "Application example 2 when combining with an emotional engine"

[1554] (Claim 1)

[1555] Means for collecting historical document data,

[1556] A means for pre-training a natural language processing model using the aforementioned past document data,

[1557] A means for fine-tuning the natural language processing model using specific parameters,

[1558] A means for receiving a request from a user and generating a draft document using the finely tuned natural language processing model,

[1559] Means for providing the generated draft document to the user,

[1560] A means for analyzing customer emotions from their statements and facial expressions, generating a draft document based on the analysis results, and displaying it on a display device,

[1561] A system that includes this.

[1562] (Claim 2)

[1563] The system according to claim 1, wherein the specific parameters for the fine-tuning learning include an amount and a time.

[1564] (Claim 3)

[1565] The system according to claim 1, wherein past document data is extracted from a document database within the company. [Explanation of Symbols]

[1566] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for collecting historical document data, A means for pre-training a natural language processing model using the aforementioned past document data, A means for fine-tuning the natural language processing model using specific parameters, A means for receiving a request from a user and generating a draft document using the finely tuned natural language processing model, Means for providing the generated draft document to the user, A system that includes this.

2. The specific parameters for the aforementioned fine-tuning learning include the amount and the timing. The system according to claim 1.

3. The historical document data was extracted from the company's internal document database. The system according to claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A